Published: 28 September 2026
NaiveAI released Naive-N0.5-Flash on 27 September 2026, an open-weight model with 309 billion parameters that switches on only 15.5 billion of them for each word. It carries an MIT licence, handles a million tokens of context, and NaiveAI prices the API at $0.10 per million input tokens. What it does not carry is a single benchmark score you can read as a number.
What is Naive-N0.5-Flash and what is it built for?
Naive-N0.5-Flash is a mixture-of-experts model, which means it switches on only part of itself for each word. NaiveAI built it on top of Xiaomi’s open-weight MiMo-V2.5 and aimed it at two jobs: writing code and running AI research experiments. The company describes the model as built with AI assistance, which is also the pitch.
The architecture is the unusual part. Most large models use full attention, where every word looks at every other word. This one has none. Its 48 layers split into 39 sliding-window layers that see only the last 128 tokens, and 9 sparse layers that pick the 2,048 most relevant tokens from anywhere in the context. That is how a model with a million-token window stays cheap to run, and it is the trade the whole design rests on.
Training ran to 3.25 trillion tokens at the full million-token context, including 50 billion tokens warming up the part that decides which tokens matter.
The practical details:
- MIT licence, so commercial use needs no negotiation
- Weights on Hugging Face at NaiveAI/Naive-N0.5-Flash, about 315 GB across 49 files
- An FP8 version and a separate 1.31 GB draft model ship alongside it
- You need FP8-capable NVIDIA GPUs, though NaiveAI does not say how many
- NaiveAI’s own server claims 50 tokens per second per user, rising to 2,000 in its fastest mode
- The API was invitation-only at launch
Naive-N0.5-Flash specifications and price
Take this from the table: the price is the headline, not the size. Naive-N0.5-Flash undercuts the cheapest comparable open model on the market by two thirds, and it is the only column where the company commits to a number.
| What is measured | Naive-N0.5-Flash | The open alternative |
|---|---|---|
| Total parameters | 309B | Xiaomi MiMo-V2.6-Pro-RL: 1.02T |
| Parameters switched on per token | 15.5B | MiMo-V2.6-Pro-RL: 42B |
| Context window | 1M tokens | MiMo-V2.6-Pro-RL: 1M tokens |
| Licence | MIT | MiMo-V2.6-Pro-RL: MIT |
| Price per million input / output tokens | $0.10 / $0.40 | MiniMax M3: $0.30 / $1.20 |
| Benchmark scores published as numbers | None | MiniMax M3: 80.5% on SWE-bench Verified |
The parameter and context figures come from NaiveAI’s model card. The MiMo and MiniMax figures come from those vendors’ own pages. NaiveAI quotes its API price in dollars on Hugging Face and in yuan on its own blog, at 0.60 and 2.60 yuan per million input and output tokens, and never puts both on one page.
What NaiveAI is not saying about its benchmarks
Every benchmark result sits inside two images. The model card names twelve tests, among them SWE-Bench Pro, DeepSWE v1.1, Terminal-Bench 2.1, MLE-bench-30 and PaperBench, and charts them against GPT-6-Astra, Opus-5.5, Gemini-3.1-Pro, GLM-5.3, Kimi-K3 and a dozen more. It prints no table. A reader who wants the numbers has to squint at a picture, and no machine can quote them at all.
The comparison is also uneven, and NaiveAI says so in small print. It ran its own model through Claude Code 2.1.207 with basic file and shell tools. The rival figures it charts were lifted from those rivals’ blogs, system cards and public leaderboards, each measured under a different setup. The card lists the sources honestly but never flags that this makes the bars on the chart non-comparable.
Missing entirely: any general-capability test, any retrieval test on the million-token window the model is sold on, and any safety evaluation. Naive-N0.5-Flash does not appear on the public SWE-bench Pro leaderboard, so nothing here has been checked by anyone outside NaiveAI.
What this means
Worth watching, not worth migrating to yet. The case for this model is entirely economic. If you run coding agents at volume, $0.10 in and $0.40 out against MiniMax M3’s $0.30 and $1.20 is the kind of gap that shows up on a monthly invoice, and an MIT licence on 309 billion parameters means nobody can take it away from you. That is a real offer to a platform team already serving open models on its own GPUs.
But you cannot justify a migration on a chart you cannot read. Until somebody publishes a number, or the model shows up on a leaderboard run by a third party, every claim here is NaiveAI’s word against nothing. If you want to test it anyway, skip the coding benchmarks and go straight at the million-token window, because that is where a design with no full attention is most likely to quietly fall over, and it is the one thing NaiveAI never measured in public.
For more information, visit the official announcement of Naive-N0.5-Flash in the NaiveAI model card.