GPTProto

OpenRouter News

Latest reporting, research and product updates filed under OpenRouter.

28 picksNewest firstLatest article Oct 1, 2026, 8:00 AM

Latest stories

3 stories
OpenRouter:Announcements(RSS)AI score 68/100

Confidence Thresholds for Model Escalation Routing

Sending every request to your strongest model gets good answers at the highest price, because you pay frontier rates for requests a cheaper model would have answered correctly. Fixed routing rules, such as picking the model by keyword or task type, cost less but need rewriting as your traffic changes. Confidence-based escalation sits between the two. You ask the model to score its own answer, then route on that score. Answers with a high score stay on a cheap model. Answers with a low score go t

OpenRouter:Announcements(RSS)AI score 67/100

How to Gate Pull Requests on LLM Evals in CI

Changing one line in a support agent’s system prompt can ship an agent that tells customers the refund window is 30 days when your policy says 14. Nothing in a normal CI pipeline checks what the model says, so the build passes and the first person to see the wrong answer is a customer.Gating a pull request on a fixed eval set works the same way as gating on a failing unit test. You keep test cases in the repository, run them when a prompt changes, and block the merge when too many fail.In this g

OpenRouter:Announcements(RSS)AI score 67/100

Cost vs. Quality Tradeoff Framework for Agent Models

Picking a model for an agent by leaderboard rank pays frontier prices for tasks that a cheaper model may handle at the same accuracy. The question to answer is not which model scores highest, but which model is the cheapest one that is good enough for the task in front of you. This guide is a three-step framework for making that call. You set the quality bar the task needs, measure cost per quality point on your own examples, and pick the cheapest model that clears the bar with margin. Tl;dr Set

3 stories
OpenRouter:Announcements(RSS)AI score 65/100

Building a Golden Eval Dataset from Production Traffic

You update a prompt, or a provider rolls out a new checkpoint under the same model ID, and something in production regresses. The last week of user complaints looks slightly different from the week before. A public benchmark like MMLU won’t catch that. It measures general capability across academic subjects, not how the model handles your product’s traffic. A golden eval dataset closes that gap. It’s a curated collection of production inputs paired with reviewed expected outputs, versioned in Gi

OpenRouter:Announcements(RSS)AI score 69/100

AI Agent Regression Testing After a Prompt or Model Change

A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail. Tl;dr Regression

OpenRouter:Announcements(RSS)AI score 62/100

How to Test Tool-Calling Accuracy in AI Agents

An agent can fail in two places when it uses a tool. It can choose the wrong tool, or it can choose the right one and send the wrong arguments. Those failures tell you different things. If an agent calls lookup_order instead of refund_order, the problem is tool selection. If it calls refund_order with the wrong order_id, it chose the right tool and passed the wrong arguments. This guide covers three ways to test tool-calling behavior. The first is a reference-free large language model (LLM) judg

1 story
1 story
1 story
1 story
1 story
1 story
Showing the latest 12 of 28 stories.