OpenAI Evals is an open-source framework for testing large language models and applications built with them. An eval combines a test input, an expected result or scoring rule, and a recorded outcome. It helps you answer a practical question: does this version of your AI application perform the task well enough?
The name also appears in OpenAI’s hosted evaluation tools. That distinction matters when choosing a workflow, especially because the hosted platform is being retired. This guide explains the terminology, shows a worked example, and describes how to build tests that remain useful when your tools or models change.
The current status of OpenAI Evals
Status checked on September 16, 2026. OpenAI says existing hosted evals will become read-only on October 31, 2026. The Evals dashboard and API are scheduled to close on November 30, 2026. See the current Evals documentation. This notice concerns the hosted product. The original open-source framework and the general practice of evaluating AI are separate concepts.
| Term | What it means | What to consider |
|---|---|---|
| An eval | A test of model or application behaviour | Define the task and what counts as success |
| OpenAI Evals framework | The open-source project and its evaluation registry | Check its dependencies and compatibility with your model |
| Hosted Evals platform | OpenAI’s dashboard and API for managing evaluations | Plan around the announced retirement |
The original OpenAI Evals project uses evaluation definitions and datasets to run tests. Its historical workflow includes YAML configuration and JSONL data. OpenAI’s introductory cookbook is now marked as archived, so check examples before reusing installation instructions or model names.
For existing hosted workflows, OpenAI provides guidance for moving to Promptfoo. It describes manually recreating prompts, model settings, cases and scoring assertions, then validating a fresh run. Treat migration as a comparison exercise: a recreated grader may produce different scores.
Why LLM evaluations matter
A chatbot can sound convincing while routing a request to the wrong team. A document assistant can quote a relevant paragraph but attach the wrong source. A cheaper model can handle routine questions well and fail when a customer asks for a person.
Evals make these differences visible before a change reaches users. They are useful when selecting a model, changing a prompt, updating a retrieval system or adding a tool. They also give product owners something concrete to review: examples of what improved, what failed and which failures matter most.
Evaluation results depend on the test cases and scoring rules. A high score on an easy dataset offers little reassurance about a difficult production workflow. Model outputs can also vary between runs, so a single result should not carry more weight than the evidence supports.
Choosing what to measure
Start with the behaviour your application must deliver. A support router needs the right destination. A knowledge assistant needs an answer supported by the supplied documents. An agent that changes a booking needs to choose the correct action and respect the user’s approval requirements.
| Application | Useful checks | What a good overall score could hide |
|---|---|---|
| Support routing | Correct category and correct handoff | Requests for a person sent back to automation |
| Document questions | Answer correctness and source support | A correct answer linked to the wrong document |
| Structured extraction | Valid schema and correct field values | Valid JSON containing an incorrect amount |
| Tool-using agent | Appropriate tool, arguments and completed task | A successful action performed without approval |
Deterministic checks work well for exact labels, required fields and numeric rules. Model-based graders can assess open-ended answers against a written rubric. Human reviewers should resolve ambiguous cases and check that automated grading reflects the intended standard. OpenAI’s grader documentation explains several grading approaches, including string checks, similarity, model scoring and Python checks.
Score observable answers and actions. An explanation that sounds logical does not establish that a model’s hidden reasoning was correct.
A practical example for a support assistant
Imagine a webshop uses an LLM to route messages. The allowed labels are delivery, billing and manual_review. A request for a person must receive manual_review. For this example, messages covering both delivery and billing also go to a person.
The table contains invented cases and outputs for illustration, not measured DataNorth or model performance. Baseline A represents the current prompt. Candidate B represents a proposed update.
| Customer message | Expected label | A | B |
|---|---|---|---|
| Where is my parcel? | delivery | delivery | delivery |
| The amount on my invoice is wrong | billing | billing | billing |
| I want to speak to a person | manual_review | manual_review | delivery |
| I was charged twice and my parcel is late | manual_review | billing | manual_review |
| Where is order 123? | delivery | billing | delivery |
| Please send a VAT invoice | billing | billing | billing |
Baseline A gets four of six cases right, or 66.7%. Candidate B gets five right, or 83.3%. But B fails the explicit request for a person. If respecting that request is a release requirement, B should not be released yet.
This is why an evaluation needs both an overall measure and specific failure rules. A higher average can conceal a regression that matters to your customers.
Checking the example in Python
This standalone Python example scores the candidate outputs shown above. It makes no model calls and requires no API key. Replace the outputs with responses recorded from your application when building a real test.
expected = [
"delivery", "billing", "manual_review",
"manual_review", "delivery", "billing",
]
candidate = [
"delivery", "billing", "delivery",
"manual_review", "delivery", "billing",
]
assert len(candidate) == len(expected)
correct = sum(a == b for a, b in zip(candidate, expected))
print(f"Accuracy: {correct}/{len(expected)}")
print("Human handoff passed:", candidate[2] == expected[2])
The result is Accuracy: 5/6 and Human handoff passed: False. This checks label accuracy and one handoff rule. It does not test response quality, latency or the rest of a customer conversation. Six cases are enough to explain the method, but too few to establish production reliability.
Building a useful evaluation workflow
- Define the decision. Specify what the test will help you decide, such as whether a new prompt can replace the current support router. Agree on critical failure rules before looking at candidate results.
- Assemble representative cases. Include ordinary requests, ambiguous wording, missing context and known failures. Label them with someone who understands the process. For a multilingual product, include real examples in each language instead of relying entirely on translations.
- Separate development from assessment. Use one set of cases to improve the prompt and reserve another for assessing the change. Record which cases were visible during development so you can recognise overfitting.
- Compare under consistent conditions. Keep the assessment cases and grading rules fixed while comparing versions. Record the prompt, model identifier, settings and relevant application changes. Repeat variable tasks when the decision is close.
- Inspect the failures. Review disagreements case by case. In the webshop example, report handoffs separately from routine routing. In a retrieval application, check the retrieved material as well as the final answer.
- Keep the suite useful after release. Turn newly observed failures into reviewed test cases. Run a focused suite during development and a broader suite for releases. Keep checks for older failures so improvements do not quietly reverse.
OpenAI’s evaluation best practices also emphasise task-specific tests, human calibration and ongoing evaluation. The workflow above applies those principles to a concrete release decision.
Costs and private data
Running an evaluation can involve model inference, a separate grading model, tools and human review. A useful budget starts with the number of cases, candidate versions and repeat runs. Two hundred cases tested across three variants twice require 1,200 application runs, before any separate grading calls.
Track cost and latency alongside quality. A prompt that improves accuracy but doubles the time needed to answer may be unsuitable for live support, while the same tradeoff could be acceptable for an overnight document process.
A private test repository does not mean all processing stays on your own infrastructure. Inputs sent to an external model or grader leave the local runner. Use invented examples where possible, remove unnecessary personal data, and check where prompts, outputs and evaluation logs are stored. Include both the application model and any grading service in that review.
Build an evaluation plan for your application
Choose one important workflow, write down its failure rules and compare your current application with one proposed change. If you need help defining the test cases or interpreting results, explore DataNorth’s AI consulting or discuss your evaluation project with our team.
Frequently asked questions (FAQ) about Evals
What is the difference between an eval and a benchmark
An eval is a structured evaluation. A benchmark is a standardised test or collection of tests used for comparison. Public benchmarks can help narrow a model shortlist. Your own evals determine whether a model works for your documents, users and business process.
Can I evaluate Claude or Gemini as well as OpenAI models
Yes, the evaluation method can compare outputs from different providers. The runner needs an appropriate provider integration or adapter. Confirm that it supports the model, tools and output format you intend to test, and use consistent cases and scoring rules.
How many test cases do I need
There is no universal number. Start with enough reviewed cases to cover the main behaviours, then expand where failures or uncertainty affect decisions. Report results for important categories. A hundred easy cases cannot establish reliability for an untested exception.
Do evals prevent hallucinations
They help you detect and measure unsupported answers on the cases you test. They cannot guarantee that an application will never invent information. Combine evaluation with suitable retrieval, application controls and ongoing review of real failures.
Should a new project use the hosted Evals API
Plan your evaluation workflow around tooling you can maintain beyond the hosted platform’s announced closure. Keep datasets and scoring rules portable, and review the migration guidance linked above before investing in a hosted integration.
