服务尚未上线Not live yet · 现在只登记意向,不收钱we are only collecting interest — nobody pays anything yet
正在筹备 · 满 80 人开机 In the works · 80 people and we launch

人人用得起的 AI AI that everyone can afford

262K 上下文,计划价 ¥4.9/月起(≈2 亿 token)——约为 Claude Opus 4.6 价格的 1%。没有 5 小时窗口,没有每周限额。服务还没上线:满 80 个想用的人,我们就开机。 A 262K context window, planned from $0.73/month (≈200M tokens) — roughly 1% of Claude Opus 4.6 pricing. No 5-hour window, no weekly cap. Not live yet: 80 people who want it and we turn it on.

frontier-class inference · ~1% of Opus 4.6 pricing · plans from $0.73, or pay as you go

现在登记不收钱 · 开机后第一批邀请发给你 · 不满意随时退出 Signing up is free · first invites go to the list · leave anytime

POST https://api.everyoneai.cc/v1/chat/completions开机后可用available at launch
# 换掉 base_url 即可,其余代码不用动
curl https://api.everyoneai.cc/v1/chat/completions \
  -H "Authorization: Bearer $EVERYONEAI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bonsai2-27b",
    "messages": [
      {"role": "system", "content": "<20 万字的系统提示词 / 知识库>"},
      {"role": "user",   "content": "总结第三章的结论"}
    ],
    "stream": true,
    "max_tokens": 2048
  }'

# 计费口径(每百万 token / per 1M tokens)
prompt_cache_hit_tokens   ¥0.01$0.002
prompt_cache_miss_tokens  ¥0.20$0.03
completion_tokens          ¥1.15$0.17
开机进度Launch progress 目标 80 人target 80
0 / 80
目标Goal
80 人80 people
已登记Signed up
0
还差Still needed
80 人80
达成后Once reached
1 周内开机live within 1 week
现在登记不收钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等——也不会收任何人的钱。 Signing up costs nothing. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to everyone on the list. Below 80, we keep waiting — and we never take anyone's money in the meantime.
套餐Token plans

一个月一杯奶茶的钱,用到够Less than a coffee a month — and it is enough

开机后的计划价格:¥4.9 / ¥9.9 / ¥19.9。额度用完后自动转按量付费——不降速、不断服。Planned pricing at launch: $0.73, $1.47 and $2.95. When the quota runs out you roll onto pay-as-you-go — no throttling, no cut-off.

尝鲜包Starter
¥4.9$0.73 / 月/ month

先试试好不好用。一顿早餐的钱。Try it out. The price of a coffee.

  • ≈ Token 总量≈ Total tokens2 亿200M
  • 计费额度Metered quota700 万单位7M units
  • ≈ Agent 调用≈ Agent calls1,800 次1,800
  • 折合Unit price¥0.700 / M$0.104 / M
  • 超出后Overflow转按量付费pay-as-you-go
标准包 · 最受欢迎Standard · most popular
¥9.9$1.47 / 月/ month

日常主力。够一个 Agent 开发者重度使用。The daily driver. Enough for an agent developer running heavy.

  • ≈ Token 总量≈ Total tokens4 亿400M
  • 计费额度Metered quota1,400 万单位14M units
  • ≈ Agent 调用≈ Agent calls3,500 次3,500
  • 折合Unit price¥0.707 / M$0.105 / M
  • 超出后Overflow转按量付费pay-as-you-go
生产力包Pro
¥19.9$2.95 / 月/ month

我们的价格上限。多项目并行、团队共享都够。Our ceiling price. Several projects or a small team.

  • ≈ Token 总量≈ Total tokens8 亿800M
  • 计费额度Metered quota2,800 万单位28M units
  • ≈ Agent 调用≈ Agent calls7,000 次7,000
  • 折合Unit price¥0.711 / M$0.105 / M
  • 超出后Overflow转按量付费pay-as-you-go
这些价格要 80 个人才能开起来These prices need 80 people to happen
现在登记不收钱,满 80 人我们开机Signing up is free — at 80 people we turn it on
登记意向Sign up

Token 量按重度 Agent 用量估算:输入:输出 ≈ 110:1,其中 97.5% 的输入命中缓存,单次调用约 11 万 token 上下文 + 1,000 token 输出。此时 1 单位 ≈ 28 个 token,一次调用约 3,970 单位。 Token estimates assume heavy agent usage: input:output ≈ 110:1, with 97.5% of input served from cache, roughly 110K tokens of context per call plus 1,000 output tokens. That is ~28 tokens per unit, or ~3,970 units per call.

对照普通对话(1.5K 输入、30% 缓存、600 输出):1 次对话约 777 单位,同样额度能跑的对话次数是上面 Agent 调用次数的 5 倍以上。缓存命中率越高、额度越耐用。 For plain chat (1.5K in, 30% cached, 600 out) a turn is ~777 units — the same quota covers more than 5× as many of those. The higher your cache-hit rate, the further your quota goes.

每月跑 1.5 万次重度 Agent 调用的用户,约需 6,000 万单位:¥19.9 档覆盖约一半,其余自动转按量付费,不会停服。 A user running 15,000 heavy agent calls a month needs roughly 60M units: the $2.95 plan covers about half, and the rest rolls onto metered billing. Nothing stops.

额度用完了怎么办?不会停服——自动转成按量付费,不降速、不排队降级。你也可以在控制台设一个月度封顶,到点自动停。 What happens when the quota runs out? Nothing stops. You roll onto pay-as-you-go — no throttling, no deprioritized queue. You can also set a monthly spend cap in the console and we stop there.

兜底费率(按量付费,每百万 token)Overflow rates (pay-as-you-go, per 1M tokens)
  • 输入 ≤64K(未命中)Input ≤64K (miss)¥0.20$0.03
  • 输入 >64K(未命中)Input >64K (miss)¥0.35$0.05
  • 输入(缓存命中)Input (cache hit)¥0.01$0.002
  • 输出(标准档)Output (Standard)¥1.15$0.17
  • 输出(极速档)Output (Turbo)¥1.90$0.28
  • 输出(异步批处理)Output (Async batch)¥0.55$0.08
对比 · Coding Planvs Coding Plans

没有 5 小时窗口,没有每周限额No 5-hour window. No weekly cap.

你不需要换客户端——Claude Code、Cline、Cursor 都能继续用,只把 base_url 换成我们的。计费方式换掉之后,那些"撞墙"就消失了。You do not have to change your client — keep using Claude Code, Cline or Cursor and just point base_url at us. Change the billing model and the walls disappear.

对比项Item 传统 Coding PlanTypical coding plan everyoneai 套餐everyoneai plans
计费窗口Billing window 5 小时滚动窗口,窗口内超额直接限流5-hour rolling window; you get throttled inside it 无窗口none
周期上限Period cap 每周限额,还分「全模型周限」和「单模型周限」两档Weekly caps — often an all-models cap plus a single-model cap 无周限,只有月度额度no weekly cap, monthly quota only
撞限之后When you hit it 窗口内只能等重置;周限撞了要等满 7 天,或者升级、转 APIWait for the session reset — or, on the weekly cap, wait 7 days, upgrade, or move to API billing 自动转按量付费,不降速不断服auto-roll to pay-as-you-go, no throttle, no cut-off
上下文Context 各家不同varies by vendor 262K 原生262K native
客户端Client 各家绑定自家客户端tied to the vendor's own client 你现有的都能用(改 base_url)keep yours — change base_url
月费对比Monthly price 对方Theirs 我们Ours 便宜Cheaper by
GitHub Copilot Pro ¥67$10 ¥4.9$0.73 13.8×
Cursor Pro ¥135$20 ¥9.9$1.47 13.6×
Claude Pro ¥135$20 ¥9.9$1.47 13.6×
Claude Max 5× ¥674$100 ¥19.9$2.95 33.9×
Claude Max 20× ¥1,348$200 ¥19.9$2.95 67.7×

对比的是计费方式与价格,不是模型能力——这些是前沿模型的订阅,我们是 27B 级开源模型。我们做的是把 Anthropic 协议一比一兼容,让你现有的客户端可以直接接过来。英文界面显示美元价(按 1 USD = 6.74 CNY 折算),对方价格以各厂商最新公告为准。This compares billing mechanics and price, not model capability — those are frontier-model subscriptions and we are a 27B-class open model. What we did is make the Anthropic protocol drop-in compatible so your existing client can point at us. USD shown in the English interface, converted at 1 USD = 6.74 CNY; vendor prices subject to their latest announcements.

~1%
相对 Opus 4.6 的输出价格of Opus 4.6 output pricing
262,144
原生上下文窗口(token)native context window (tokens)
84.4 / 100
1,179 题公开基准池实测on a 1,179-item public benchmark pool
Apache 2.0
权重开源,可自行部署open weights, self-hostable
能力Capabilities

不是"便宜的小模型",是"能读完你整份文档的模型"Not a cheap small model — a model that can read your entire document

27B 参数、三元量化、原生 262K 窗口。为长文档、RAG、代码库问答和批量处理而设计。27B parameters, ternary-quantized, native 262K window. Built for long documents, RAG, codebase Q&A and batch processing.

◧

262K 原生窗口Native 262K context

不是 RoPE 外推。单次请求可读入约 20 万字中文文档,检索位置实测覆盖到 96 万 token。Not RoPE extrapolation. Read roughly 200k Chinese characters — or a 600-page PDF — in a single request; retrieval verified up to 960k tokens.

◈

27B 级质量,5.9GB 权重27B-class quality, 5.9 GB of weights

三元 {-1, 0, +1} 权重配 FP16 分组缩放,保留全精度 Qwen3.8-27B 的 98.2% 能力,推理成本降到十分之一。Ternary {-1, 0, +1} weights with FP16 group scales retain 98.2% of the full-precision Qwen3.8-27B, at a tenth of the serving cost.

◱

前缀缓存按命中计费Prefix caching, billed on hits

系统提示词、知识库、多轮历史只算一次。缓存命中 ¥0.01/M,只有未命中的二十分之一。System prompts, knowledge bases and conversation history are charged once. Cache hits cost $0.002/M — one twentieth of a miss.

⇄

双协议零改造Drop-in for two protocols

同时提供 OpenAI Chat Completions 与 Anthropic Messages 接口。现有 SDK 与 Agent 框架改 base_url 就能用。Both OpenAI Chat Completions and Anthropic Messages endpoints. Point your existing SDK or agent framework at a new base_url and you are done.

◐

视觉 · 工具调用 · 思考模式Vision · tools · thinking mode

支持图片与多图输入、function calling、流式输出;思考过程可保留,也可关闭以降低延迟与成本。Image and multi-image input, function calling, streaming. Keep or disable the reasoning trace to trade latency and cost.

◍

三档服务,按场景选Three tiers, pick your trade-off

标准档走批量并发;极速档独占低并发通道,单路 200+ tok/s;异步批处理最便宜,¥0.55/M,24 小时内交付。Standard runs batched. Turbo gets a dedicated low-concurrency lane at 200+ tok/s per stream. Async batch is cheapest at $0.08/M, delivered within 24 hours.

对比Comparison

同样是 262K 级的活儿,价格差 100 倍The same 262K-class job, at 1/100th of the price

左边是当今最强的前沿模型之一 Claude Opus 4.6,右边是我们。差别写在明面上,你自己判断该用哪个。On the left, one of today's strongest frontier models, Claude Opus 4.6. On the right, us. The difference is stated plainly — you decide which one your task needs.

每百万 tokenPer 1M tokens Claude Opus 4.6 everyoneai.cc 便宜倍数Cheaper by
输入(未命中)Input (cache miss) ¥33.70$5.00 ¥0.20$0.03 135×
输入(缓存命中)Input (cache hit) ¥3.37$0.50 ¥0.01$0.002 112×
输出Output ¥168.50$25.00 ¥1.15$0.17 105×
上下文窗口Context window 1,000,000 262,144 0.26×0.26×
模型规模Model scale 前沿闭源Frontier, closed 27B 开源27B, open weights —

买套餐的话,输出成本可低至 ¥0.70 / 百万单位(¥19.9 生产力包),相当于 Opus 4.6 的 0.42%。On the $2.95 Pro plan, output-equivalent cost drops to $0.10 per 1M units — about 0.42% of Opus 4.6.

Opus 4.6 价格取自 Anthropic 2026-02 公开发布价,汇率按 1 USD = 6.74 CNY。以对方最新公告为准。Opus 4.6 pricing as published by Anthropic (Feb 2026); converted at 1 USD = 6.74 CNY. Subject to their latest published rates.

按量付费Pay as you go

套餐额度用完之后的兜底价The rate you roll onto after the quota

不买套餐也能直接用。三档服务按速度定价,单路 70–200 tok/s,输出 ¥0.55–1.90 / 百万 token。买套餐的话,同样的量再便宜 12–22%。No plan required. Three tiers priced by speed — 70 to 200 tok/s per stream, $0.08–0.28 per 1M output tokens. With a plan the same volume costs another 12–22% less.

服务档Tier 并发额度Concurrency 单路输出速度Speed per stream 输出价格 ¥/MOutput $/M 适用场景Best for
极速档Turbo 独占 ≤2 路dedicated ≤2 130–200 tok/s ¥1.90$0.28 交互式编程、实时客服live coding, real-time support
标准档Standard 共享 16 路池shared 16-lane pool 70–100 tok/s ¥1.15$0.17 对话、RAG、内容生成chat, RAG, content generation
异步批处理Async batch 队列调度queued 不承诺延迟no latency SLA ¥0.55$0.08 离线清洗、批量摘要、评测offline cleaning, bulk summarization, evals

速度为实测值(单路、短上下文、开启投机解码)。长上下文(>128K)下单路输出约 90 tok/s;聚合吞吐随并发线性提升,当前服务能力约 1,100 tok/s。Speeds are measured single-stream with speculative decoding on short context. Beyond 128K context expect ~90 tok/s per stream. Aggregate throughput scales with concurrency; current service capacity is about 1,100 tok/s.

计费项Line item 换算成单位Units 说明Notes
输出 token(标准档)Output tokens (Standard)1 token = 1 单位基准单位the base unit
输入 token ≤64K(未命中)Input ≤64K (miss)1 token = 0.16 单位首次读入的提示词first read of a prompt
输入 token >64K(未命中)Input >64K (miss)1 token = 0.35 单位长文档一次性读入one-shot long-document read
输入 token(缓存命中)Input (cache hit)1 token = 0.02 单位相同前缀再次出现,几乎免费same prefix again — nearly free
输出 token(异步批处理)Output (async batch)1 token = 0.5 单位24 小时内交付,半价delivered within 24h, half price
按量价目Metered rates 人民币 / 百万 tokenUSD / 1M tokens 说明Notes
输入(≤64K,未命中)Input (≤64K, miss)¥0.20$0.03首次读入first read
输入(>64K,未命中)Input (>64K, miss)¥0.35$0.05长文档一次性读入long-document read
输入(缓存命中)Input (cache hit)¥0.01$0.002相同前缀再次出现same prefix again
输出(标准档)Output (Standard)¥1.15$0.17批量并发batched
输出(极速档)Output (Turbo)¥1.90$0.28独占低并发通道,单路 200+ tok/sdedicated lane, 200+ tok/s per stream
输出(异步批处理)Output (Async batch)¥0.55$0.0824 小时内交付delivered within 24h

英文界面显示美元价,按 1 USD = 6.74 CNY 折算并取整;人民币价格为准。USD prices are shown in the English interface, converted at 1 USD = 6.74 CNY and rounded; CNY is the reference price.

接入Get started

三行代码迁移Three lines to migrate

兼容 OpenAI 与 Anthropic 两套协议。已有代码只需替换 base_url 与 API Key。Compatible with both protocols. Replace the base_url and the API key — that is the whole migration.

Python · OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://api.everyoneai.cc/v1",
    api_key="$EVERYONEAI_KEY",
)

resp = client.chat.completions.create(
    model="bonsai2-27b",
    messages=[{"role": "user",
               "content": "总结这份合同的风险点"}],
    stream=True,
)
Python · Anthropic SDK
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.everyoneai.cc",
    api_key="$EVERYONEAI_KEY",
)

msg = client.messages.create(
    model="bonsai2-27b",
    max_tokens=2048,
    messages=[{"role": "user",
               "content": "总结这份合同的风险点"}],
)
常见问题FAQ

你可能想知道的Things you probably want to know

现在能用吗?Can I use it today?

还不能。机器一台没买,现在只登记意向,不收任何人的钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等。Not yet. No hardware has been bought and we take no money at this stage — we are only collecting interest. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to the list. Below 80, we keep waiting.

为什么是 80?因为这个规模能让我们把机器开起来、并撑住前三个月的运营。达不到就是白烧钱,我不想那样开始。Why 80? Because that is the scale at which the machine can be switched on and kept running for the first three months. Below it, we would just be burning money — and I would rather not start that way.

262K 上下文是真的吗?还是只是参数上限?Is the 262K context real, or just a parameter?

是模型原生窗口,不是外推。我们已在 262,144 token 的文档上做过针检索测试:埋在全文 33%、66%、90% 位置的三段信息全部按序取回。上游引擎的实测上限更高(在该硬件上可开到 66 万至 97 万 token)。It is the model's native window, not extrapolation. We ran a needle test on a 262,144-token document: codes planted at 33%, 66% and 90% were all retrieved in order. The upstream engine has been measured even higher on this hardware — 660k to 970k tokens.

这是什么模型?和 Qwen3.8-27B 是什么关系?What model is behind this, and how does it relate to Qwen3.8-27B?

Ternary Bonsai 2 27B,基于 Qwen3.8-27B 架构以三元权重原生训练({-1, 0, +1}),配 FP16 分组缩放,权重约 5.9GB,官方基准保留全精度模型 98.2% 的综合能力。权重以 Apache 2.0 开源,你也可以自行部署。Ternary Bonsai 2 27B — natively trained with ternary weights ({-1, 0, +1}) on the Qwen3.8-27B architecture, with FP16 group scales. About 5.9 GB of weights, retaining 98.2% of the full-precision model's aggregate benchmark score. Released under Apache 2.0, so you can self-host it too.

为什么能比 Opus 4.6 便宜 100 倍?质量差多少?How can this be 100× cheaper than Opus 4.6? How much worse is it?

便宜来自三件事:模型只有 7GB 出头(一张消费级显卡就能跑)、三元权重让显存带宽瓶颈大幅缓解、我们用批量并发把单卡利用率拉满。质量上,它在公开基准池拿 84.4/100,做分类、抽取、摘要、翻译、客服问答这类任务和前沿模型差距很小;但复杂推理、长链路 Agent、难题编程明显不如 Opus 4.6。Three things: the model is just over 7 GB (it runs on a single consumer GPU), ternary weights relieve the memory-bandwidth bottleneck, and we batch requests to keep utilization high. On quality it scores 84.4/100 on a public benchmark pool — close to frontier models on classification, extraction, summarization, translation and support Q&A, but clearly behind Opus 4.6 on complex reasoning, long-horizon agents and hard coding.

缓存命中怎么判定、怎么计费?How are cache hits determined and billed?

服务端按前缀精确匹配复用 KV 缓存。响应 usage 中区分 prompt_cache_hit_tokens 与 prompt_cache_miss_tokens,分别按 ¥0.01/M 与 ¥0.20/M 计费。固定系统提示词、同一份长文档的多轮追问、Agent 的重复上下文都能命中。We do exact prefix matching on the server and reuse the KV cache. The usage block separates prompt_cache_hit_tokens from prompt_cache_miss_tokens, billed at $0.002/M and $0.20/M respectively. Fixed system prompts, follow-up questions on the same long document and repeated agent context all hit.

容量用完了会怎样?扩容要多久?What happens when capacity runs out, and how long does scaling take?

我们会按需增加 GPU 实例,但新增实例需要「拉起机器 → 拉取 8.87GiB 模型权重 → 预热」,约 5–10 分钟。满载时新请求会进入队列并按提交顺序处理,不会丢失。如果你有可预期的峰值,提前告知我们会事先扩容;紧急情况也可以通过控制台申请优先通道。We add GPU instances on demand, but a new instance has to boot, pull the 8.87 GiB model artifact and warm up — roughly 5–10 minutes. When saturated, new requests queue in submission order and are never dropped. Tell us about expected peaks and we scale ahead of time; you can also request a priority lane from the console.

套餐额度怎么算?输入 token 也算吗?How is plan quota counted — do input tokens count too?

算,但按成本折算成"单位"。1 个输出 token = 1 单位;输入 token(≤64K)只算 0.16 单位;长文档输入(>64K)算 0.35 单位;缓存命中只要 0.02 单位。所以把系统提示词和长文档交给缓存,额度几乎不掉。完整换算表在上面的"按量付费"一节。Yes, but converted into units by cost. 1 output token = 1 unit; input tokens under 64K cost 0.16 units; long-document input above 64K costs 0.35 units; cache hits cost just 0.02 units. Reuse your system prompt and long documents through the cache and your quota barely moves. Full table is in the pay-as-you-go section above.

可以随时退订吗?用不完的额度能留到下个月吗?Can I cancel? Does unused quota roll over?

随时可退,按未使用天数比例退款。当月未用完的额度可结转一个月(第二个月底清零)。额度用完后不会停服,会自动转按量付费,你也可以在控制台设一个月度封顶自动停。Cancel anytime with a pro-rata refund for unused days. Unused quota rolls over for one month and expires at the end of the following month. When the quota runs out we do not stop you — you roll onto pay-as-you-go, and you can set a monthly spend cap in the console.

和 Claude Code 那种订阅有什么区别?How is this different from a Claude Code style subscription?

三点:没有 5 小时滚动窗口,没有每周限额,价格便宜一个数量级。传统订阅撞到周限后只能等满 7 天、升级或转 API;我们额度用完后自动转按量付费,不降速、不断服。而且我们不绑定客户端——Claude Code、Cline、Cursor 改个 base_url 就能接到我们这里。Three things: no 5-hour rolling window, no weekly cap, and an order of magnitude cheaper. On a classic subscription, hitting the weekly cap means waiting seven days, upgrading, or switching to API billing; with us the quota simply rolls into pay-as-you-go with no throttle and no cut-off. And we are not tied to a client — point Claude Code, Cline or Cursor at our base_url and you are done.

捐款真的会让我的额度变多吗?Will donating actually raise my quota?

会。捐款的 95% 进入额度池,用于增加算力和统一上调三档额度——上调对全体用户生效,包括没有捐款的用户,也包括你。剩下的 5% 用于支付通道手续费与账务成本。每月 1 号公示上月收到的金额、进入额度池的金额,以及额度实际提升了多少。额度池按累积制计算,累计达标后 30 天内完成上调。Yes. 95% of every donation goes into the quota pool, which is spent on adding capacity and raising all three plan quotas — the increase applies to every user, including those who never donated, and including you. The remaining 5% covers payment processing and accounting. On the 1st of each month we publish the total received, the amount that went into the pool, and how much the quotas actually rose. The pool is cumulative: increases land within 30 days of crossing a threshold.

捐款不产生任何额外权益:不会给你更多额度、不会插队、不会解锁功能。额度池的余额和流水在控制台随时可查。这不是慈善募捐,是用户自愿多付、用来把价格压给所有人的赞助。Donating grants no extra privileges: no bonus quota, no queue priority, no unlocked features. The pool balance and full ledger are visible in the console at any time. Not a charity appeal — it is users voluntarily paying a little extra, spent on keeping the price low for everyone.

我的数据会被用于训练吗?Is my data used for training?

不会。请求内容仅用于本次推理,不用于训练、不对外共享。企业版可提供数据不出境与完整私有化部署。No. Request content is used only to serve that request — never for training, never shared. An enterprise tier offers data residency and fully private deployment.

人人用得起的 AI,需要 80 个人一起开始AI everyone can afford — starting with 80 people

现在登记不收钱,也不绑卡。满 80 人我们开机,第一批邀请发给你。Signing up costs nothing and needs no card. At 80 people we launch, and the first invites go out.

登记意向Sign up 看计划价格See planned pricing