Salesforce launches Koa, a CRM reasoning model

16-09-2026

Salesforce and NVIDIA announced Koa on 15 September 2026, Salesforce's first CRM reasoning model for Agentforce. Koa is a post-trained NVIDIA Nemotron-3-Super-120B and scores 69.41 on Tau2Bench, ahead of GPT-4.1 but 14.58 points behind GPT-5.5.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

salesforce launches koa, a crm reasoning model built on nvidia's 120b nemotron
Sign up for our Newsletter

Published on: 16 September 2026

Salesforce and NVIDIA announced Koa on 15 September 2026, Salesforce’s first reasoning model for its Agentforce agent platform. Koa is a post-trained version of NVIDIA’s open-weight Nemotron-3-Super-120B, and Salesforce’s own technical report scores it at 69.41 on the Tau2Bench customer service benchmark, ahead of GPT-4.1 but 14.58 points behind GPT-5.5. It is in pilot with a handful of customers now, with general availability expected in winter 2026 in United States regions only.

What is Salesforce Koa?

Koa is a language model tuned to do one job: work through a multi-step CRM task and call the right tools along the way. Think of qualifying a lead, routing a service case or scheduling a follow-up. Each of those needs several actions in the right order rather than one good answer.

Salesforce built it by post-training Nemotron-3-Super-120B, NVIDIA’s 120-billion-parameter open-weights model, using supervised fine-tuning and reinforcement learning with Group Relative Policy Optimization, a method that scores a batch of attempts against each other. The tooling came from NVIDIA: NeMo RL, NeMo Gym and NeMo AutoModel. The reinforcement learning run used five NVIDIA B200 nodes.

No customer data went into training. The corpus is synthetic: scenarios written to mirror workflows across more than 14 industries, each pairing a persona with a task and the sequence of tool calls needed to finish it.

The commercial point is control. Salesforce holds the weights, does the post-training and runs inference inside its own infrastructure, so no customer data crosses the trust boundary at inference time. You cannot download Koa and you cannot call it outside Agentforce. Salesforce already runs it internally as a Slack agent, and pilot customers include 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine and Xero.

Salesforce Koa benchmarks: how it compares to GPT-5.5 and Claude Opus 4.8

The figures below come from Salesforce’s own technical report, posted to arXiv on 14 September. Read them as a model that beats its own starting point by roughly a point and loses to the current frontier by ten or more.

BenchmarkSalesforce KoaBest score in Salesforce’s own table
Tau2Bench (multi-turn customer service)69.4183.99 (OpenAI GPT-5.5)
BFCL (calling tools correctly)66.6378.18 (Claude Opus 4.8)
CRM Bench (Salesforce workflows, 0 to 1)0.860.90 (OpenAI GPT-5.5)
CRM function-call accuracy (0 to 1)0.770.85 (OpenAI GPT-4.1)
Tau2Bench, against its own base model69.4168.64 (Nemotron-3-Super-120B)
BFCL, against its own base model66.6364.73 (Nemotron-3-Super-120B)

Every figure here is Salesforce’s own. The paper says it compared Koa with three proprietary models, but it names no API version, date or setting for any of them, so the rival column is not a controlled head-to-head. Salesforce’s own checkpoints were served through vLLM at BF16 precision.

What Salesforce is not saying

The press release says Koa “matches or exceeds leading model performance on CRM actions with three times fewer errors” and links the paper as the source. That phrase does not appear in the paper. No table in the paper reports an error rate at all. The paper’s own one-line summary reads differently: Koa “surpasses a strong proprietary baseline while remaining below the strongest frontier models”. The baseline in question is GPT-4.1, which OpenAI shipped in April 2025.

Salesforce also publishes an ablation that undercuts its own shipping choice. A checkpoint trained with supervised fine-tuning alone scores 70.04 on Tau2Bench and 0.88 on CRM Bench, against 69.41 and 0.86 for the reinforcement-learning checkpoint that became Koa. Salesforce picked the lower-scoring one because it is much better at multi-turn tool use, where it reaches 59.50 on BFCL against 53.25. That is a defensible trade, and it is more candid than the press release.

Three numbers are missing entirely: price, context window and latency. For a model sold as a hosted option inside Agentforce, the price is the one that decides whether anyone tries it.

What this means

Koa is worth watching rather than planning around, and the interesting part is the method, not the model. Salesforce took an open-weights 120B model, spent a modest amount of compute on reinforcement learning driven by written workflow specifications, and closed the gap to GPT-4.1 while keeping every byte of inference inside its own perimeter. Any large enterprise sitting on a decade of process documentation and answering to a regulator should read that paper, because the recipe generalises further than the model does.

The model itself is a narrower proposition. If you run Agentforce, sit in a United States region, and your blocker is that customer data cannot leave your vendor’s perimeter, put your hand up for the pilot. Everyone else can wait. Koa trails GPT-5.5 by 14.58 points on the benchmark Salesforce chose to lead with, there is no published price, and general availability is a season away. A specialised model that sits below the general-purpose frontier has to win on cost, on control, or on both. So far Salesforce has only published the control half.

For more information, visit the official announcement of Koa in the Salesforce newsroom.

Add DataNorth AI to your Google favorites