ori

Ori Eval: # Find the best model for what you're building

tell your agent
❯
# Using OpenRouter MCP? Run /spawn-ori-eval in your agent to get started
read the docs →

01──Find the best model, with proof

Just ask. Your agent will hand the request to Ori Eval, which explores your codebase, runs evaluations on models, and comes back with an answer.

02──New to evals? We figure out what to test with you

Never written an eval? Ori Eval scans your codebase for every place a model runs — which surface, where the code is, what model it uses today — and checks which one you want evals on. It guides you like an experienced engineering friend.

03──Consistent results you can trust, run after run

Keeping the test bench identical, run after run, makes evaluations repeatable and consistent. That's why Ori Eval is an agent: unlike a skill, it's able to pin model and effort. It comes pre-tuned - we've already chosen the harness and model that works best, so you don't have to.

04──Catch a bug before you fix it

Every bug becomes a test: it fails now, passes once you fix the agent, and keeps the bug from coming back.

05──Stop regressions before they ship

In CI, a bad change fails the build, so bugs never reach your users.

06──Re-run as new models drop

Your evals are just code. Manually re-run them when a new model drops, or schedule them monthly. You can sleep well knowing you always have the best model for what you're building.

07──Know it did the right thing

An eval checks three things: the tools the agent called, the tools it avoided, and the quality of the answer — so you catch the wrong move before your users do.

08──faq

More and more apps use AI models as part of their core functionality, but the
choice of which model to use is often made without a systematic method. There
is no definitive “best model”, only the best model for what you’re building.
Ori Eval helps you find that model, with quantitative proof.

Ori Eval is for anyone that wants to choose the best model for what they're
building through quantitative proof rather than guesswork. You do not need to
be an evaluation expert; we built it so anyone can use it. It interviews and
guides you through the entire model evaluation process, like a really smart
engineering coworker.

One of the harder things about running model evaluations is keeping the test
harness consistent across runs - you don't want multiple runs of the same
valuation to give you different results. As a result, Ori Eval's agent harness
and model are pinned, so every run scores the same setup. After gathering
requirements with you, Ori Eval builds *.eval.ts files and runs them with bun
test.

Ori pins the harness and model during a run. The environment does not change,
so if there are differences, it would be caused by the model.

The majority of harnesses that people use are limited to the company's AI
models. Ori Eval has access to all of OpenRouter's models, which gives Ori
Eval the flexibility to use models with the best price-to-performance ratio.
We handle the test harness for you, so you can focus on improving your
project.

Yes. For interactive agent use, run ori login. Ori opens OpenRouter in your
browser and saves the credential for later runs, so you do not create or paste
an API key. If no credential is available, Ori stops before the model call and
tells you to sign in. For CI, use OPENROUTER_API_KEY instead.

For CI, set OPENROUTER_API_KEY from a repository secret. Ori does not need ori
login when this variable is set.

Ori Eval will get a list of the latest models on OpenRouter that fit your
criteria, and tell you why it was chosen. You don't need to get the list
yourself; you don't even need to know anything about AI models. The model
selection process is highly customizable -- you can set a price ceiling, keep
your current model in the mix, or run the same eval against each one.

Yes, it is highly encouraged, in fact -- the more data you feed it about your
use case, the more accurate the evaluation run gets. You will be asked by Ori
Eval whether you want your data included in the evaluations. The eval runs
every row and grades a new model against your current one on your own gold
answers.

We use LLM-as-a-judge to grade open-ended answers. Ori Eval helps establish a
grading criteria and a minimum score.

Add ori eval to a GitHub Actions workflow. It exits like bun test, so a failed
eval fails the build. Make it a separate opt-in job — evals call real models
and cost money.

Create a *.eval.ts file, import setupAgent from ori/eval, call
agent.run(prompt), and assert on the run: run.tool("x").toBeCalled(),
run.toComplete(), run.toMention("…"). Then run ori eval.

You usually don't have to: tell your coding agent to run curl -fsSL
https://openrouter.ai/skills/spawn-ori-eval and follow the instructions in its
output to get started, and the skill installs Ori and runs the eval for you.
To install by hand, run: curl -fsSL https://openrouter.ai/labs/ori/install.sh
| bash

No. The OpenRouter MCP has a spawn-ori-eval tool. Run /spawn-ori-eval in your
agent to get started — same document, no URL to paste.
tell your agent❯