262K 原生窗口Native 262K context
不是 RoPE 外推。单次请求可读入约 20 万字中文文档,检索位置实测覆盖到 96 万 token。Not RoPE extrapolation. Read roughly 200k Chinese characters — or a 600-page PDF — in a single request; retrieval verified up to 960k tokens.
262K 上下文,计划价 ¥4.9/月起(≈2 亿 token)——约为 Claude Opus 4.6 价格的 1%。没有 5 小时窗口,没有每周限额。服务还没上线:满 80 个想用的人,我们就开机。 A 262K context window, planned from $0.73/month (≈200M tokens) — roughly 1% of Claude Opus 4.6 pricing. No 5-hour window, no weekly cap. Not live yet: 80 people who want it and we turn it on.
frontier-class inference · ~1% of Opus 4.6 pricing · plans from $0.73, or pay as you go
现在登记不收钱 · 开机后第一批邀请发给你 · 不满意随时退出 Signing up is free · first invites go to the list · leave anytime
# 换掉 base_url 即可,其余代码不用动
curl https://api.everyoneai.cc/v1/chat/completions \
-H "Authorization: Bearer $EVERYONEAI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bonsai2-27b",
"messages": [
{"role": "system", "content": "<20 万字的系统提示词 / 知识库>"},
{"role": "user", "content": "总结第三章的结论"}
],
"stream": true,
"max_tokens": 2048
}'
# 计费口径(每百万 token / per 1M tokens)
prompt_cache_hit_tokens ¥0.01$0.002
prompt_cache_miss_tokens ¥0.20$0.03
completion_tokens ¥1.15$0.17
开机后的计划价格:¥4.9 / ¥9.9 / ¥19.9。额度用完后自动转按量付费——不降速、不断服。Planned pricing at launch: $0.73, $1.47 and $2.95. When the quota runs out you roll onto pay-as-you-go — no throttling, no cut-off.
先试试好不好用。一顿早餐的钱。Try it out. The price of a coffee.
日常主力。够一个 Agent 开发者重度使用。The daily driver. Enough for an agent developer running heavy.
我们的价格上限。多项目并行、团队共享都够。Our ceiling price. Several projects or a small team.
Token 量按重度 Agent 用量估算:输入:输出 ≈ 110:1,其中 97.5% 的输入命中缓存,单次调用约 11 万 token 上下文 + 1,000 token 输出。此时 1 单位 ≈ 28 个 token,一次调用约 3,970 单位。 Token estimates assume heavy agent usage: input:output ≈ 110:1, with 97.5% of input served from cache, roughly 110K tokens of context per call plus 1,000 output tokens. That is ~28 tokens per unit, or ~3,970 units per call.
对照普通对话(1.5K 输入、30% 缓存、600 输出):1 次对话约 777 单位,同样额度能跑的对话次数是上面 Agent 调用次数的 5 倍以上。缓存命中率越高、额度越耐用。 For plain chat (1.5K in, 30% cached, 600 out) a turn is ~777 units — the same quota covers more than 5× as many of those. The higher your cache-hit rate, the further your quota goes.
每月跑 1.5 万次重度 Agent 调用的用户,约需 6,000 万单位:¥19.9 档覆盖约一半,其余自动转按量付费,不会停服。 A user running 15,000 heavy agent calls a month needs roughly 60M units: the $2.95 plan covers about half, and the rest rolls onto metered billing. Nothing stops.
额度用完了怎么办?不会停服——自动转成按量付费,不降速、不排队降级。你也可以在控制台设一个月度封顶,到点自动停。 What happens when the quota runs out? Nothing stops. You roll onto pay-as-you-go — no throttling, no deprioritized queue. You can also set a monthly spend cap in the console and we stop there.
你的捐款 95% 会进入「额度池」,用于提升所有用户的套餐额度——包括你自己。剩下 5% 用于支付通道手续费与账务成本。95% of your donation goes into the quota pool and is used to raise the plan quota for every user — including you. The other 5% covers payment processing and accounting.
额度池怎么用:只做两件事——① 把三档套餐额度统一上调;② 增加算力,让上调后的额度真的用得上。没有第三件事。 What the pool is used for: exactly two things — (1) raising all three plan quotas together; (2) adding capacity so the higher quota can actually be used. Nothing else.
每月公示:1 号公布上月收到的捐款总额、进入额度池的金额,以及三档额度实际提升了多少。捐款不产生任何额外权益——额度提升对全体用户一视同仁。
额度什么时候生效:额度池按累积制计算。累计金额达到门槛后,我们会在接下来的 30 天内完成扩容与额度上调,并在公示页写明具体生效日期。不是达标即刻生效——扩容需要时间,所以我们统一在月度节点上调,避免中途反复改规则。 Published monthly: on the 1st we post the total received, the amount that went into the pool, and how much the quotas actually rose. Donating grants no extra privileges — the increase applies to everyone equally.
When the increase lands: the pool is cumulative. Once the running total crosses a threshold we complete the GPU purchase and raise the quotas within the next 30 days, and we publish the exact effective date. It is not instant — capacity takes time to bring online — so increases are applied at a monthly boundary rather than changing the rules mid-cycle.
捐款要等开机后才开放。现在登记不收任何钱。开机后额度池的余额和流水会在控制台公开。Donations open only after launch. Signing up costs nothing today. Once we are live, this is how the pool works: users voluntarily paying a little extra, spent entirely on raising everyone's quota. The pool balance and ledger are visible in the console at any time.
你不需要换客户端——Claude Code、Cline、Cursor 都能继续用,只把 base_url 换成我们的。计费方式换掉之后,那些"撞墙"就消失了。You do not have to change your client — keep using Claude Code, Cline or Cursor and just point base_url at us. Change the billing model and the walls disappear.
| 对比项Item | 传统 Coding PlanTypical coding plan | everyoneai 套餐everyoneai plans |
|---|---|---|
| 计费窗口Billing window | 5 小时滚动窗口,窗口内超额直接限流5-hour rolling window; you get throttled inside it | 无窗口none |
| 周期上限Period cap | 每周限额,还分「全模型周限」和「单模型周限」两档Weekly caps — often an all-models cap plus a single-model cap | 无周限,只有月度额度no weekly cap, monthly quota only |
| 撞限之后When you hit it | 窗口内只能等重置;周限撞了要等满 7 天,或者升级、转 APIWait for the session reset — or, on the weekly cap, wait 7 days, upgrade, or move to API billing | 自动转按量付费,不降速不断服auto-roll to pay-as-you-go, no throttle, no cut-off |
| 上下文Context | 各家不同varies by vendor | 262K 原生262K native |
| 客户端Client | 各家绑定自家客户端tied to the vendor's own client | 你现有的都能用(改 base_url)keep yours — change base_url |
| 月费对比Monthly price | 对方Theirs | 我们Ours | 便宜Cheaper by |
|---|---|---|---|
| GitHub Copilot Pro | ¥67$10 | ¥4.9$0.73 | 13.8× |
| Cursor Pro | ¥135$20 | ¥9.9$1.47 | 13.6× |
| Claude Pro | ¥135$20 | ¥9.9$1.47 | 13.6× |
| Claude Max 5× | ¥674$100 | ¥19.9$2.95 | 33.9× |
| Claude Max 20× | ¥1,348$200 | ¥19.9$2.95 | 67.7× |
对比的是计费方式与价格,不是模型能力——这些是前沿模型的订阅,我们是 27B 级开源模型。我们做的是把 Anthropic 协议一比一兼容,让你现有的客户端可以直接接过来。英文界面显示美元价(按 1 USD = 6.74 CNY 折算),对方价格以各厂商最新公告为准。This compares billing mechanics and price, not model capability — those are frontier-model subscriptions and we are a 27B-class open model. What we did is make the Anthropic protocol drop-in compatible so your existing client can point at us. USD shown in the English interface, converted at 1 USD = 6.74 CNY; vendor prices subject to their latest announcements.
27B 参数、三元量化、原生 262K 窗口。为长文档、RAG、代码库问答和批量处理而设计。27B parameters, ternary-quantized, native 262K window. Built for long documents, RAG, codebase Q&A and batch processing.
不是 RoPE 外推。单次请求可读入约 20 万字中文文档,检索位置实测覆盖到 96 万 token。Not RoPE extrapolation. Read roughly 200k Chinese characters — or a 600-page PDF — in a single request; retrieval verified up to 960k tokens.
三元 {-1, 0, +1} 权重配 FP16 分组缩放,保留全精度 Qwen3.8-27B 的 98.2% 能力,推理成本降到十分之一。Ternary {-1, 0, +1} weights with FP16 group scales retain 98.2% of the full-precision Qwen3.8-27B, at a tenth of the serving cost.
系统提示词、知识库、多轮历史只算一次。缓存命中 ¥0.01/M,只有未命中的二十分之一。System prompts, knowledge bases and conversation history are charged once. Cache hits cost $0.002/M — one twentieth of a miss.
同时提供 OpenAI Chat Completions 与 Anthropic Messages 接口。现有 SDK 与 Agent 框架改 base_url 就能用。Both OpenAI Chat Completions and Anthropic Messages endpoints. Point your existing SDK or agent framework at a new base_url and you are done.
支持图片与多图输入、function calling、流式输出;思考过程可保留,也可关闭以降低延迟与成本。Image and multi-image input, function calling, streaming. Keep or disable the reasoning trace to trade latency and cost.
标准档走批量并发;极速档独占低并发通道,单路 200+ tok/s;异步批处理最便宜,¥0.55/M,24 小时内交付。Standard runs batched. Turbo gets a dedicated low-concurrency lane at 200+ tok/s per stream. Async batch is cheapest at $0.08/M, delivered within 24 hours.
左边是当今最强的前沿模型之一 Claude Opus 4.6,右边是我们。差别写在明面上,你自己判断该用哪个。On the left, one of today's strongest frontier models, Claude Opus 4.6. On the right, us. The difference is stated plainly — you decide which one your task needs.
| 每百万 tokenPer 1M tokens | Claude Opus 4.6 | everyoneai.cc | 便宜倍数Cheaper by |
|---|---|---|---|
| 输入(未命中)Input (cache miss) | ¥33.70$5.00 | ¥0.20$0.03 | 135× |
| 输入(缓存命中)Input (cache hit) | ¥3.37$0.50 | ¥0.01$0.002 | 112× |
| 输出Output | ¥168.50$25.00 | ¥1.15$0.17 | 105× |
| 上下文窗口Context window | 1,000,000 | 262,144 | 0.26×0.26× |
| 模型规模Model scale | 前沿闭源Frontier, closed | 27B 开源27B, open weights | — |
买套餐的话,输出成本可低至 ¥0.70 / 百万单位(¥19.9 生产力包),相当于 Opus 4.6 的 0.42%。On the $2.95 Pro plan, output-equivalent cost drops to $0.10 per 1M units — about 0.42% of Opus 4.6.
Opus 4.6 价格取自 Anthropic 2026-02 公开发布价,汇率按 1 USD = 6.74 CNY。以对方最新公告为准。Opus 4.6 pricing as published by Anthropic (Feb 2026); converted at 1 USD = 6.74 CNY. Subject to their latest published rates.
不买套餐也能直接用。三档服务按速度定价,单路 70–200 tok/s,输出 ¥0.55–1.90 / 百万 token。买套餐的话,同样的量再便宜 12–22%。No plan required. Three tiers priced by speed — 70 to 200 tok/s per stream, $0.08–0.28 per 1M output tokens. With a plan the same volume costs another 12–22% less.
| 服务档Tier | 并发额度Concurrency | 单路输出速度Speed per stream | 输出价格 ¥/MOutput $/M | 适用场景Best for |
|---|---|---|---|---|
| 极速档Turbo | 独占 ≤2 路dedicated ≤2 | 130–200 tok/s | ¥1.90$0.28 | 交互式编程、实时客服live coding, real-time support |
| 标准档Standard | 共享 16 路池shared 16-lane pool | 70–100 tok/s | ¥1.15$0.17 | 对话、RAG、内容生成chat, RAG, content generation |
| 异步批处理Async batch | 队列调度queued | 不承诺延迟no latency SLA | ¥0.55$0.08 | 离线清洗、批量摘要、评测offline cleaning, bulk summarization, evals |
速度为实测值(单路、短上下文、开启投机解码)。长上下文(>128K)下单路输出约 90 tok/s;聚合吞吐随并发线性提升,当前服务能力约 1,100 tok/s。Speeds are measured single-stream with speculative decoding on short context. Beyond 128K context expect ~90 tok/s per stream. Aggregate throughput scales with concurrency; current service capacity is about 1,100 tok/s.
| 计费项Line item | 换算成单位Units | 说明Notes |
|---|---|---|
| 输出 token(标准档)Output tokens (Standard) | 1 token = 1 单位 | 基准单位the base unit |
| 输入 token ≤64K(未命中)Input ≤64K (miss) | 1 token = 0.16 单位 | 首次读入的提示词first read of a prompt |
| 输入 token >64K(未命中)Input >64K (miss) | 1 token = 0.35 单位 | 长文档一次性读入one-shot long-document read |
| 输入 token(缓存命中)Input (cache hit) | 1 token = 0.02 单位 | 相同前缀再次出现,几乎免费same prefix again — nearly free |
| 输出 token(异步批处理)Output (async batch) | 1 token = 0.5 单位 | 24 小时内交付,半价delivered within 24h, half price |
| 按量价目Metered rates | 人民币 / 百万 tokenUSD / 1M tokens | 说明Notes |
|---|---|---|
| 输入(≤64K,未命中)Input (≤64K, miss) | ¥0.20$0.03 | 首次读入first read |
| 输入(>64K,未命中)Input (>64K, miss) | ¥0.35$0.05 | 长文档一次性读入long-document read |
| 输入(缓存命中)Input (cache hit) | ¥0.01$0.002 | 相同前缀再次出现same prefix again |
| 输出(标准档)Output (Standard) | ¥1.15$0.17 | 批量并发batched |
| 输出(极速档)Output (Turbo) | ¥1.90$0.28 | 独占低并发通道,单路 200+ tok/sdedicated lane, 200+ tok/s per stream |
| 输出(异步批处理)Output (Async batch) | ¥0.55$0.08 | 24 小时内交付delivered within 24h |
英文界面显示美元价,按 1 USD = 6.74 CNY 折算并取整;人民币价格为准。USD prices are shown in the English interface, converted at 1 USD = 6.74 CNY and rounded; CNY is the reference price.
兼容 OpenAI 与 Anthropic 两套协议。已有代码只需替换 base_url 与 API Key。Compatible with both protocols. Replace the base_url and the API key — that is the whole migration.
from openai import OpenAI
client = OpenAI(
base_url="https://api.everyoneai.cc/v1",
api_key="$EVERYONEAI_KEY",
)
resp = client.chat.completions.create(
model="bonsai2-27b",
messages=[{"role": "user",
"content": "总结这份合同的风险点"}],
stream=True,
)
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.everyoneai.cc",
api_key="$EVERYONEAI_KEY",
)
msg = client.messages.create(
model="bonsai2-27b",
max_tokens=2048,
messages=[{"role": "user",
"content": "总结这份合同的风险点"}],
)
还不能。机器一台没买,现在只登记意向,不收任何人的钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等。Not yet. No hardware has been bought and we take no money at this stage — we are only collecting interest. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to the list. Below 80, we keep waiting.
为什么是 80?因为这个规模能让我们把机器开起来、并撑住前三个月的运营。达不到就是白烧钱,我不想那样开始。Why 80? Because that is the scale at which the machine can be switched on and kept running for the first three months. Below it, we would just be burning money — and I would rather not start that way.
是模型原生窗口,不是外推。我们已在 262,144 token 的文档上做过针检索测试:埋在全文 33%、66%、90% 位置的三段信息全部按序取回。上游引擎的实测上限更高(在该硬件上可开到 66 万至 97 万 token)。It is the model's native window, not extrapolation. We ran a needle test on a 262,144-token document: codes planted at 33%, 66% and 90% were all retrieved in order. The upstream engine has been measured even higher on this hardware — 660k to 970k tokens.
Ternary Bonsai 2 27B,基于 Qwen3.8-27B 架构以三元权重原生训练({-1, 0, +1}),配 FP16 分组缩放,权重约 5.9GB,官方基准保留全精度模型 98.2% 的综合能力。权重以 Apache 2.0 开源,你也可以自行部署。Ternary Bonsai 2 27B — natively trained with ternary weights ({-1, 0, +1}) on the Qwen3.8-27B architecture, with FP16 group scales. About 5.9 GB of weights, retaining 98.2% of the full-precision model's aggregate benchmark score. Released under Apache 2.0, so you can self-host it too.
便宜来自三件事:模型只有 7GB 出头(一张消费级显卡就能跑)、三元权重让显存带宽瓶颈大幅缓解、我们用批量并发把单卡利用率拉满。质量上,它在公开基准池拿 84.4/100,做分类、抽取、摘要、翻译、客服问答这类任务和前沿模型差距很小;但复杂推理、长链路 Agent、难题编程明显不如 Opus 4.6。Three things: the model is just over 7 GB (it runs on a single consumer GPU), ternary weights relieve the memory-bandwidth bottleneck, and we batch requests to keep utilization high. On quality it scores 84.4/100 on a public benchmark pool — close to frontier models on classification, extraction, summarization, translation and support Q&A, but clearly behind Opus 4.6 on complex reasoning, long-horizon agents and hard coding.
服务端按前缀精确匹配复用 KV 缓存。响应 usage 中区分 prompt_cache_hit_tokens 与 prompt_cache_miss_tokens,分别按 ¥0.01/M 与 ¥0.20/M 计费。固定系统提示词、同一份长文档的多轮追问、Agent 的重复上下文都能命中。We do exact prefix matching on the server and reuse the KV cache. The usage block separates prompt_cache_hit_tokens from prompt_cache_miss_tokens, billed at $0.002/M and $0.20/M respectively. Fixed system prompts, follow-up questions on the same long document and repeated agent context all hit.
我们会按需增加 GPU 实例,但新增实例需要「拉起机器 → 拉取 8.87GiB 模型权重 → 预热」,约 5–10 分钟。满载时新请求会进入队列并按提交顺序处理,不会丢失。如果你有可预期的峰值,提前告知我们会事先扩容;紧急情况也可以通过控制台申请优先通道。We add GPU instances on demand, but a new instance has to boot, pull the 8.87 GiB model artifact and warm up — roughly 5–10 minutes. When saturated, new requests queue in submission order and are never dropped. Tell us about expected peaks and we scale ahead of time; you can also request a priority lane from the console.
算,但按成本折算成"单位"。1 个输出 token = 1 单位;输入 token(≤64K)只算 0.16 单位;长文档输入(>64K)算 0.35 单位;缓存命中只要 0.02 单位。所以把系统提示词和长文档交给缓存,额度几乎不掉。完整换算表在上面的"按量付费"一节。Yes, but converted into units by cost. 1 output token = 1 unit; input tokens under 64K cost 0.16 units; long-document input above 64K costs 0.35 units; cache hits cost just 0.02 units. Reuse your system prompt and long documents through the cache and your quota barely moves. Full table is in the pay-as-you-go section above.
随时可退,按未使用天数比例退款。当月未用完的额度可结转一个月(第二个月底清零)。额度用完后不会停服,会自动转按量付费,你也可以在控制台设一个月度封顶自动停。Cancel anytime with a pro-rata refund for unused days. Unused quota rolls over for one month and expires at the end of the following month. When the quota runs out we do not stop you — you roll onto pay-as-you-go, and you can set a monthly spend cap in the console.
三点:没有 5 小时滚动窗口,没有每周限额,价格便宜一个数量级。传统订阅撞到周限后只能等满 7 天、升级或转 API;我们额度用完后自动转按量付费,不降速、不断服。而且我们不绑定客户端——Claude Code、Cline、Cursor 改个 base_url 就能接到我们这里。Three things: no 5-hour rolling window, no weekly cap, and an order of magnitude cheaper. On a classic subscription, hitting the weekly cap means waiting seven days, upgrading, or switching to API billing; with us the quota simply rolls into pay-as-you-go with no throttle and no cut-off. And we are not tied to a client — point Claude Code, Cline or Cursor at our base_url and you are done.
会。捐款的 95% 进入额度池,用于增加算力和统一上调三档额度——上调对全体用户生效,包括没有捐款的用户,也包括你。剩下的 5% 用于支付通道手续费与账务成本。每月 1 号公示上月收到的金额、进入额度池的金额,以及额度实际提升了多少。额度池按累积制计算,累计达标后 30 天内完成上调。Yes. 95% of every donation goes into the quota pool, which is spent on adding capacity and raising all three plan quotas — the increase applies to every user, including those who never donated, and including you. The remaining 5% covers payment processing and accounting. On the 1st of each month we publish the total received, the amount that went into the pool, and how much the quotas actually rose. The pool is cumulative: increases land within 30 days of crossing a threshold.
捐款不产生任何额外权益:不会给你更多额度、不会插队、不会解锁功能。额度池的余额和流水在控制台随时可查。这不是慈善募捐,是用户自愿多付、用来把价格压给所有人的赞助。Donating grants no extra privileges: no bonus quota, no queue priority, no unlocked features. The pool balance and full ledger are visible in the console at any time. Not a charity appeal — it is users voluntarily paying a little extra, spent on keeping the price low for everyone.
不会。请求内容仅用于本次推理,不用于训练、不对外共享。企业版可提供数据不出境与完整私有化部署。No. Request content is used only to serve that request — never for training, never shared. An enterprise tier offers data residency and fully private deployment.
现在登记不收钱,也不绑卡。满 80 人我们开机,第一批邀请发给你。Signing up costs nothing and needs no card. At 80 people we launch, and the first invites go out.