Articles written by Jorick van Weelie
Cloudflare releases Clef and Clef-flash
Cloudflare released Clef and Clef-flash on 1 October 2026, two open-weight decision models that return probabilities instead of text. Clef-flash decides in 38.8 milliseconds at the median, both carry a
Microsoft releases MAI-Transcribe-2-Streaming
Microsoft released MAI-Transcribe-2-Streaming on 1 October 2026, its first real-time transcription model, with MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft reports a 2.5% word error rate and a final transcript 0.13 seconds after
Perplexity releases pplx-embed-v2-context-9b-preview
Perplexity released pplx-embed-v2-context-9b-preview on 30 September 2026, an MIT-licensed embedding model with open weights on Hugging Face.
Liquid AI releases d1
Liquid AI released d1 on 29 September 2026, a decision model that classifies, routes and scores without generating a single output token.
ElevenLabs releases Eleven v4 and Eleven v4 Turbo
ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026. Turbo starts speaking after about 150 milliseconds, against 262 for Cartesia Sonic 3.6 and 814 for OpenAI's GPT-4o
MiniMax releases M3.1-Flash-Preview
MiniMax released M3.1-Flash-Preview on 27 September 2026, a coding model with a 1,000,000-token context that runs only inside the company's MiniMax Code tool.
Jev: What it is and why this new model matters
Jev is TypeSafe AI’s first public System One model. Instead of generating text it evaluates a state against typed questions and returns choices, scores and probabilities with a confidence estimate.
Liquid AI released LFM2.5-VL-3B-DSpark
Liquid AI's new DSpark drafter accelerates LFM2.5-VL-3B without replacing the underlying vision model. Vendor tests show up to 3.13x faster decoding and 2.62x end-to-end gains, with the largest benefit where
Google releases Gemini 3.8 Flash TTS
Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. Flash-Lite costs $6 per million audio output tokens until the end of the year, then
Recraft releases Recraft V4.1 Flash
Recraft released Recraft V4.1 Flash on 23 September 2026, generating a roughly 1,000 pixel image in about 1.5 seconds for $0.007, around a fifth of the price of Recraft V4.1.
Anthropic releases Claude Opus 5.5
Anthropic released Claude Opus 5.5 on 22 September 2026 at $4 per million input tokens and $20 per million output tokens, which it says is 40 percent cheaper to run
What it costs to keep an AI system running after
Running an AI system costs four things: cloud and model usage, maintenance hours, ownership and coordination, and retraining where a trained model is involved. Infrastructure is routinely overestimated and maintenance
Dutch AI factory picks Zernike in Groningen
On 21 September 2026 SURF announced that the Dutch AI factory's supercomputer will be housed at Eurofiber's datacenter campus on the Zernike Campus in Groningen. The machine holds about 1,800
xAI adds Grok Voice Transcribe 2.0
xAI has added Grok Voice Transcribe 2.0 to its batch and streaming speech-to-text API while keeping version 1.0 as the default. Published list prices remain $0.10 per hour for REST
Alibaba Qwen3.8 adds 1M-context Omni analysis and 60-language live translation
Alibaba Cloud has added two Qwen3.8 multimodal endpoints. Omni-Flash combines a 1M context window with audio, video, reasoning and tools, while LiveTranslate targets real-time multilingual translation with speech output.
PrismML releases Ternary Bonsai 2 27B
On 17 September 2026 PrismML released Ternary Bonsai 2 27B, a version of Alibaba's Qwen3.8 27B. The model shrinks from 54 GB to 5.9 GB and keeps 98.2 percent of
Deepgram released Nova-3 Pharma
Deepgram's Nova-3 Pharma is a specialist speech-to-text model for drug names and pharmaceutical vocabulary. It works in batch and streaming through the hosted API and can be requested for self-hosting.
NVIDIA releases DeepSeek-V4.1-Flash-NVFP4
NVIDIA published DeepSeek-V4.1-Flash-NVFP4 on 16 September 2026, a four-bit conversion of DeepSeek-V4.1-Flash for Blackwell GPUs. It runs on four GB300 GPUs under MIT licence, supports 1 million tokens of context,
Creatify launches Boreal video model
Creatify Labs' Boreal turns LTX-2.5 into an advertising-focused video model priced from one cent per generated second at 720p. That economics can materially increase creative iteration, although the launch benchmark
Salesforce launches Koa, a CRM reasoning model
Salesforce and NVIDIA announced Koa on 15 September 2026, Salesforce's first CRM reasoning model for Agentforce. Koa is a post-trained NVIDIA Nemotron-3-Super-120B and scores 69.41 on Tau2Bench, ahead of GPT-4.1