GPTProto

Deployment & Ops News

Latest reporting, research and product updates filed under Deployment & Ops.

555 picksNewest firstLatest article Oct 1, 2026, 8:00 AM

Latest stories

7 stories
OpenRouter:Announcements(RSS)AI score 68/100

Confidence Thresholds for Model Escalation Routing

Sending every request to your strongest model gets good answers at the highest price, because you pay frontier rates for requests a cheaper model would have answered correctly. Fixed routing rules, such as picking the model by keyword or task type, cost less but need rewriting as your traffic changes. Confidence-based escalation sits between the two. You ask the model to score its own answer, then route on that score. Answers with a high score stay on a cheap model. Answers with a low score go t

Modal 官方工程博客(RSS)AI score 61/100

VM Sandboxes: Full computers for agents

Today we’re making VM Sandboxes generally available on Modal, built for those who need to give their agents the power of a full computer.With one flag, you’ll get a fully capable Linux VM with all the niceties that you expect from a traditional modal.Sandbox, and it Just Works™. This brings the same APIs, modal.Images, sub-second cold-starts, and CPU/memory bursting capabilities as previously, all whilst supporting the hundreds of thousands of concurrent Sandboxes that our users are accustomed t

OpenRouter:Announcements(RSS)AI score 67/100

How to Gate Pull Requests on LLM Evals in CI

Changing one line in a support agent’s system prompt can ship an agent that tells customers the refund window is 30 days when your policy says 14. Nothing in a normal CI pipeline checks what the model says, so the build passes and the first person to see the wrong answer is a customer.Gating a pull request on a fixed eval set works the same way as gating on a failing unit test. You keep test cases in the repository, run them when a prompt changes, and block the merge when too many fail.In this g

Modal 官方工程博客(RSS)AI score 63/100

Runtime Roundup: VM Sandboxes, Multi-node clusters, and more

Modal just hosted our inaugural conference, Runtime. Here are a few of the highlights that we announced.VM SandboxesVM Sandboxes give your agent access to a full Linux computer. Agents increasingly want to live inside something that looks like a real machine: running Docker stacks, local databases and dev servers, graphical environments and mobile simulators, and even monkeying around with the Linux Kernel itself.VM Sandboxes are already in use at customers like Linear, Legora, and Snorkel power

OpenRouter:Announcements(RSS)AI score 67/100

Cost vs. Quality Tradeoff Framework for Agent Models

Picking a model for an agent by leaderboard rank pays frontier prices for tasks that a cheaper model may handle at the same accuracy. The question to answer is not which model scores highest, but which model is the cheapest one that is good enough for the task in front of you. This guide is a three-step framework for making that call. You set the quality bar the task needs, measure cost per quality point on your own examples, and pick the cheapest model that clears the bar with margin. Tl;dr Set

Google DeepMind:Blog(RSS)AI score 78/100

Gemini 4 Argon: our next era of frontier intelligence

Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. In this article Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Go

3 stories
OpenRouter:Announcements(RSS)AI score 65/100

Building a Golden Eval Dataset from Production Traffic

You update a prompt, or a provider rolls out a new checkpoint under the same model ID, and something in production regresses. The last week of user complaints looks slightly different from the week before. A public benchmark like MMLU won’t catch that. It measures general capability across academic subjects, not how the model handles your product’s traffic. A golden eval dataset closes that gap. It’s a curated collection of production inputs paired with reviewed expected outputs, versioned in Gi

OpenRouter:Announcements(RSS)AI score 69/100

AI Agent Regression Testing After a Prompt or Model Change

A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail. Tl;dr Regression

OpenAI:官网动态(RSS · 排除企业/客户案例)AI score 86/100

Introducing GPT-6.1 Sol

Near-Astra intelligence for a fifth of the price We’re introducing GPT‑6.1 Sol, an upgrade to GPT‑6 Sol that nearly matches GPT‑6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing—giving developers more room to build and run capable agents that reuse context across request

2 stories
3 stories
Anthropic:Claude.dev 开发者博客(RSS)AI score 67/100

Automating eval design and hillclimbing with Claude

Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill. With the skill, you can run /claude-api build-eval to build an evaluation inside your codebase, and run /claude-api hillclimb to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting. In this article, we

Hacker News:AI 热帖AI score 86/100

Prompting Claude Opus 5.5

This guide covers the prompting patterns specific to Claude Opus 5.5. For the model's capabilities and API changes, see What's new in Claude Opus 5.5. For techniques that apply across all current Claude models, see Prompting best practices. Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and tends to finish the same task with fewer tokens. Existing Claude Opus 5 prompts should perform well without changes, and the patterns in Prompting Claude Opus 5 remain

Tomer Tunguz 博客(VC 分析)AI score 63/100

How GPU Prices Can Double While AI Gets Cheaper

GPU prices have doubled in the last six months, from $4.40 to $8.08 per GPU-hour. But AI prices are falling. How can that be? It is not a simple answer. Higher GPU costs are likely to remain while the industry races to build out new data centers. Every component of the buildout is increasing in cost, from concrete to copper to credit.1 Above all, electricity remains the limiting factor, which Oracle is experiencing : last week it invoked force majeure on its New Mexico campus after the natural-g

1 story
1 story
2 stories
3 stories
1 story
2 stories
1 story
1 story
2 stories
1 story
3 stories
1 story
4 stories
2 stories
1 story
2 stories
1 story
1 story
1 story
1 story
2 stories
2 stories
2 stories
1 story
1 story
Showing the latest 55 of 555 stories.