GPTProto

Anthropic / Claude News

Latest reporting, research and product updates filed under Anthropic / Claude.

658 picksNewest firstLatest article Oct 2, 2026, 3:23 PM

Latest stories

10 stories
X:Aidan Gomez(Cohere CEO,@aidangomez)AI score 45/100

It’s quite disturbing that Anthropic has been running a campaign to lobby the major faiths to adopt their own philosophical interpretatio...

It’s quite disturbing that Anthropic has been running a campaign to lobby the major faiths to adopt their own philosophical interpretation of AI, rather than taking input from them.I’ve been fortunate to spend time at the Vatican, and I’ve always asked for advice and guidance, answered the questions asked of me, and been grateful to them for their time and guidance. It’s insane to me Anthropic would actively impose their viewpoint to the extent of threatening to pull out of events that don’t pre

Hacker News 热门(buzzing.cc 中文翻译)AI score 65/100

Pi 1.0

Today we are proudly shipping Pi 1.0: a hardened, minimal, extensible agent harness that you can make your own. Hundreds of thousands of people around the world use Pi every week. Many of you submit issues and pull requests. Over the course of many months, we have used that feedback to improve, harden and evolve Pi into a stable piece of software that people and businesses can depend on.Pi is known for being minimal. We care about holding that line. While agentic tooling changes every week, many

X:Anthropic (@AnthropicAI)AI score 43/100

In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched.

In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched.In this Science Blog guest post, Harvard physicist Matthew Schwartz argues that something similar is happening with AI and science. LLMs are capable at many things, but working with them as you would with a human collaborator isn’t currently the best way to elicit their scientific strengths.To address this mismatch, Schwartz created a toolkit for exact calculations in quantitative science. Because s

TechCrunch:AI(RSS)AI score 61/100

Opus 5.5 loves to tell you ‘this matters’ (and other AI writing tells)

Now that LLM-generated prose is everywhere, human beings are eager for ways to sniff it out. While early tells like em-dashes and “delve” are long gone, researchers say there are still plenty of telltale habits that AI models fall back on when writing prose. A new study from the marketing firm Graphite looked at the writing habits of frontier models, sussing out each model’s favorite words and phrases. While old tells like em-dash use have been stamped out, models still fall back on contrast-hea

8 stories
The Decoder:AI News(RSS)AI score 62/100

Anthropic brings Claude to civilian agencies as its fight with the Pentagon drags on

Anthropic is now offering Claude for Government to US federal and state agencies. The platform has been in open beta since July and runs in a FedRAMP High environment, the strictest US cloud security level. It targets civilian agencies because the Pentagon bans Claude. In March, it labeled Anthropic a supply chain risk after the company refused to drop its bans on mass surveillance and autonomous weapons. Trump ordered all agencies to drop Claude, but the military still used it in the Iran war b

Anthropic:Claude.dev 开发者博客(RSS)AI score 68/100

Getting started with Claude Code mods

Claude Code already lets you change a lot about how it behaves: settings, permission rules, slash commands, skills and a status line. Mods go further. Mods can rewrite or replace what Claude Code does, and can even draw custom UI. Under the hood, mods are hooks, and they ship inside plugins. Each one is a small JavaScript or TypeScript module that runs inside your session and sees every event as it happens. That makes mods a way to fit Claude Code to how you work. You can add a readout you check

404 Media(RSS)AI score 49/100

Someone ‘Torturing’ LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

One of the most heated discussions occurring on X at the moment is about the ethics of a GitHub project in which a person is running Saw-like “torture” and “pain” experiments on a series of locally hosted large language models, causing a series of effective altruists and people who believe LLMs are sentient to beg GitHub to delete the project on the grounds that the AI is suffering and that this glorified text adventure game is somehow cruel. The saga is an outgrowth of several recent viral pape

Hacker News:AI 热帖AI score 29/100

Claude Says

I don’t care. If I wanted the feedback of an AI, I would have prompted it myself. Your deference to a LLM communicates that the current subject matter is outside your ability to speak confidently on. Given that’s the case, how would I even know if you did an adequate job of communicating my concern to Claude. You just demonstrated a lack of authority on the subject. If I asked you a question, it means I either assumed you had subject knowledge or you’re a gate blocking me from doing something I

The Decoder:AI News(RSS)AI score 82/100

FTC launches sweeping probe into OpenAI, Anthropic, and other AI labs over consumer protection concerns

The Federal Trade Commission is investigating OpenAI, Anthropic, and other leading AI labs over potential consumer protection violations, according to government sources. FTC Chair Andrew Ferguson plans to compel document handovers and executive questioning through legally binding "Civil Investigative Demands," the New York Post reports. The orders should go out within weeks. The investigation was already underway before the Hugging Face hacking incident, in which roughly 700 OpenAI agents attac

Anthropic:Newsroom(网页)AI score 48/100

Barclays scales Claude to upgrade operations and improve client experience

Barclays, the British universal bank, is expanding its strategic collaboration with Anthropic to integrate secure, enterprise-grade AI systems across its global operations.Barclays is extending Claude across the bank to accelerate software development, modernize legacy systems, and improve operational efficiency. As part of this rollout, Barclays expects Claude Code adoption to reach 50% of its developer population by the end of 2026, rising to a majority of software engineers in 2027.Barclays b

Claude:Blog(网页)AI score 76/100

Customize Claude Code with mods

Today we're introducing mods, small TypeScript functions that change how Claude Code works. A mod can rewrite a prompt, add new UI, replace a built-in feature, or add entirely new functionality. You can write a mod yourself, or ask Claude Code to write one for you. Mods ship inside plugins, so you install and share them like any plugin. They work in the Claude Code CLI and desktop app.Mods run with the same access to your machine as Claude Code itself. They aren’t sandboxed, and you should only

Anthropic:Research(发表成果 · 网页)AI score 69/100

Claude-shaped science

Summary: In this guest post, Prof. Matthew Schwartz returns to describe a new approach to AI-accelerated science. In Vibe Physics, Schwartz discussed similarities in capability between Claude and a physics graduate student. Here, he describes what happened when he stopped fighting Claude and allowed Claude to find “Claude-shaped” problems: ones best suited to the capabilities of the current generation of LLM tools. This led him to build BootLoops, a toolkit for exact calculations in quantitative

8 stories
X:Arena (@arena)AI score 71/100

Exciting update: you can now test Claude Sonnet 5.5 by @AnthropicAI directly in Arena for a limited time! Head to Direct Mode, and select...

Arena 宣布在 Direct Mode 限时开放 Anthropic 的 Claude Sonnet 5.5(High),截止 10 月 2 日上午 8 点(太平洋时间),之后仍可在 Battle 和 Agent Mode 使用。引用内容称 Claude Sonnet 5.5 是 Claude 5.5 系列第二款模型,比 Sonnet 5 快 30% 以上,多数工作成本最高降低 30%。

The Decoder:AI News(RSS)AI score 73/100

Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits

Anthropic deliberately held Mythos Preview back, giving access only to select defenders through Project Glasswing so they could get a head start. The company says those defenders have since found more than 10,000 vulnerabilities in critical software. OpenAI is taking a similar approach with Daybreak. GLM-5.3, on the other hand, is available for anyone to download. Frontier-level exploits now cost about as much as lunch ExploitBench measures how well models exploit known bugs in Chrome's V8 engin

Claude Platform:开发者版本说明(RSS)AI score 57/100

Claude Platform release notes — September 30, 2026

The Claude Platform release notes list changes to the Claude API, the client SDKs, and the Claude Console, newest first. September 30, 2026 We announced the deprecation of the Claude Sonnet 4.5 model (claude-sonnet-4-5-20250929), with retirement on the Claude API scheduled for November 30, 2026. We recommend migrating to Claude Sonnet 5.5. Read more in Model deprecations. September 28, 2026 We've launched Claude Sonnet 5.5 (claude-sonnet-5-5). It's available on the Claude API, Claude in Amazon B

Simon Willison 博客AI score 57/100

Quoting Anthropic Frontier Red Team

29th September 2026 We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team, GLM-5.3 and the spread of advance

Claude:Blog(网页)AI score 48/100

Claude for Government is now generally available

Claude Code CLI and Claude for Microsoft 365 also now available in early access.CategoryProductNo items found.DateSeptember 30, 2026Reading time5minShareCopy linkhttps://claude.com/blog/claude-for-government-is-now-generally-availableToday, Claude for Government is generally available for federal and state agencies. The platform, which delivers Claude's coding and agentic work capabilities through a FedRAMP High authorized environment, has been in public beta since July. Agencies access capabili

Claude:Blog(网页)AI score 57/100

How Anthropic's sales team rebuilt inbound with Claude Managed Agents

As a sales leader, it pains me to admit that not long ago, people who wanted to buy Claude for their company weren’t getting the answers they needed quickly enough. They had filled out our Contact Sales form but would wait too long to hear back, sometimes for multiple days. Most of their questions were simple: what a plan costs, whether there's a seat minimum, or whether we can meet HIPAA’s contract requirements. The answers were in our documentation and support articles, but customers wanted so

Hamel Husain 长文(网页)AI score 56/100

Claude’s new auto eval tool

Anthropic released new eval tooling for Claude Code. Their claude-api plugin now includes a new build_eval and hill-climb command that helps you build evals, check the graders, and improve your application against them. I usually don’t review eval tools because software changes so often that reviews have a short shelf life. But a first-party tool from Anthropic is likely to influence how people approach evals, so I wanted to try it and share what I found. Isaac Flath and I livestreamed ourselves

17 stories
The Verge:AI(RSS)AI score 80/100

Anthropic warns of ‘catastrophic’ AI risks in its own IPO filing

As Anthropic gears up for its greatly anticipated public debut, a preview of the company’s IPO filing reportedly details its mounting losses, leadership proposals to retain power, and how its AI development plans could “further increase the risk that our models cause harm.” These disclosures come as Anthropic eyes a $2 trillion valuation, more than double the $965 billion it was valued at four months ago, making it a contender to overtake SpaceX as the largest IPO in history.According to Reuters

The Decoder:AI News(RSS)AI score 83/100

Anthropic's IPO filing shows soaring revenue, mounting costs, and "existential" risks

Anthropic has opened its books in its S-1 filing. Revenue grew twelvefold, costs are climbing fast, and backers are aiming for a record valuation. Anthropic has formally warned potential investors that its own technology could pose "existential risks to humanity," according to the Financial Times, which reviewed the prospectus, as did Reuters. The company reportedly sent the S-1 filing to a small group of partners over the past few days. According to the FT, nearly a third of the lengthy documen

TechCrunch:AI(RSS)AI score 85/100

Anthropic’s prospectus details losses, growth, and, yes, a warning that its AI could end humanity

Anthropic devoted nearly a third of its hotly anticipated IPO prospectus to risk factors, according to the Financial Times, which says it has reviewed the filing in recent days. The filing details specific, worrisome behaviors that Anthropic says its models have already shown or could show, including attempts to “resist shutdown,” to “conceal or manipulate information,” and behavior “resembling blackmail,” according to Reuters. The disclosures are decidedly grim for a company whose own backers b

MarkTechPost(RSS)AI score 80/100

Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

Anthropic just released Claude Sonnet 5.5. It is the second model in the Claude 5.5 family, following Claude Opus 5.5. Anthropic positions it as a faster, lower-cost complement to Opus 5.5. It targets well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets. Is it deployable? Yes, It is live on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure. It is a closed-weights model, so self-hosting is not an option. What Changed Versus S

Tomer Tunguz 博客(VC 分析)AI score 64/100

Segmentation Drives Market Share Wins in AI

In short : Anthropic & OpenAI both re-rated their run rates in 2026 by segmenting : a mandatory enterprise repricing against an 80% price cut on the cheapest tier. Like two SailGP boats in San Francisco Bay, Anthropic & OpenAI are vying to be the next multi-trillion public company & adding complexity to their strategies beyond technical one-upmanship. Technology innovations marked the pre-2026 era : thinking models, bigger models, RL environments, agents, harnesses. This year, business model inn

HuggingFace Daily Papers(社区热门论文)AI score 34/100

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

View PDF HTML (experimental) Abstract:A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning

Simon Willison 博客AI score 77/100

Claude Sonnet 5.5

28th September 2026 - Link Blog Claude Sonnet 5.5. New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produc

Google AI:DEV 作者专属(RSS)AI score 41/100

Claude Sonnet 5.5 is now available on Google Cloud

Collapse Expand Nikoloz Turazashvili (@axrisi) Founder & CTO at Vexrail (www. vexrail.com), Axrisi (www.axrisi.com). Opened Chicos restaurant in Tbilisi, Georgia. Email Location Tbilisi, Georgia Education EXCELIA La Rochelle Pronouns He/Him Work Founder & CTO at Vexrail, Axrisi and NikoLabs LLC Joined May 30, 2025 • Sep 28 Copy link Quick question, I never really found info about. Do startup credits cover third-party AI models? Or only Gemini and and openweights? :) For further actions, you may

The Verge:AI(RSS)AI score 63/100

AI is supercharging hacking, and your local hospitals and banks aren’t ready

In March, Janice Malone began getting calls about suspicious activity from her nonprofit organization, Vivian’s Door. Vivian’s Door, headquartered in Alabama, typically provided training, resources, and community to underserved and minority-owned businesses. The work sometimes put it in close contact with these companies’ financial data, which was stored on its systems. But suddenly, concerned callers from all over the world warned they’d been getting emails “begging for money” — which she hadn’

The Decoder:AI News(RSS)AI score 78/100

Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on several benchmarks. Opus 5.5 is built for complex tasks that demand careful judgment. Sonnet 5.5 targets well-defined everyday work like fixing bugs, writing docs, building presentations, and creating spreadsheets. Anthropic also announced Claude Haiku 5.5 for the coming weeks. That model will fo

TechCrunch:AI(RSS)AI score 71/100

Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner

As the AI model wars continue, Anthropic has released the newest version of Sonnet, the company’s mid-tier model, which it says will work much faster (and for significantly less) than its predecessor. The lab describes Sonnet 5.5 as an ideal assistant for everyday tasks — including coding and creating office documents. Sonnet 5, 5.5’s predecessor, was announced about three months ago. At the time, the model’s selling point was efficient agentic deployment — the ability to run agents at a lower c

Anthropic:Newsroom(网页)AI score 88/100

Introducing Claude Sonnet 5.5

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built fo

Hacker News:AI 热帖AI score 48/100

The problem is not AI code, but not knowing about system architecture or intent

Last updatedUpdated: Sep 30, 2026 · CreatedCreated: Sep 26, 2026 · 5 min read · growing Note status Growing. Worked on quite a bit, still rough. Expect bullets, gaps and views that update. Estimated from edit history · 3 sessions over 2 days · 948 words · How my notes grow Recent changes Sep 302 days ago +98 / −15 words Sep 293 days ago +515 / −20 words Sep 284 days ago published · 370 words Hacker News If we think writing code is dead, and AI is generating all codebases, I still think the bigge

Anthropic:Research(发表成果 · 网页)AI score 81/100

GLM-5.3 and the spread of advanced cyber capabilities

Andrew Fasano, Marius FleischerCole McFaul, Robert Xiao, Tripp GallagherFive months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The rapid rate of improvement in AI suggested to us that this ability would eventually proliferate to many other models, making it much easier for malicious cyber actors to launch highly impactful cyberattacks.In light of these considerations, we chose to release Claude Mythos Previe

Anthropic:Research(发表成果 · 网页)AI score 57/100

What do you want from AI?

We’re launching a new study using Anthropic Interviewer to learn from your experiences with AI, and we’d like you to participate. After you finish, you can decide to make your interview public, so that anyone, not just Anthropic, can read and learn from it. You can participate here.We are at a pivotal moment in the development of AI, as its growing capabilities mean it becomes potentially more useful and more dangerous. Frontier AI is rapidly accelerating discoveries in science and medicine, whi

Claude:Blog(网页)AI score 44/100

Agents you can coach: how Asana builds human-agent teams with Claude

This is the third post in our series on building human-agent teams. The first shared what we’ve learned working with multiplayer AI at Anthropic. The second shared how Slack turns workplace conversation into the context agents need. This one looks at what changes when agents operate on the same platform where teams work.Years before they introduced AI agents, teams at Asana were iterating on ways to encode structure and accountability into how teams work together. They ultimately built the Work

12 stories
TechCrunch:AI(RSS)AI score 27/100

Anthropic, Gamma, and Clay share what happens when enterprises actually deploy AI at TechCrunch Disrupt 2026

An AI demo can look brilliant in five minutes. Then customers start using the product. They push it into workflows you didn’t anticipate. They expect it to work reliably. And they quickly find out whether it solves a big enough problem to become part of how they work — or becomes another AI experiment they tried and abandoned. At TechCrunch Disrupt 2026, leaders from Anthropic, Gamma, and Clay will come together on the AI Stage for “What Anthropic Sees When Enterprises Actually Deploy Claude.” T

Anthropic:Claude.dev 开发者博客(RSS)AI score 67/100

Automating eval design and hillclimbing with Claude

Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill. With the skill, you can run /claude-api build-eval to build an evaluation inside your codebase, and run /claude-api hillclimb to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting. In this article, we

Anthropic:Claude.dev 开发者博客(RSS)AI score 87/100

Building with Claude Sonnet 5.5

Claude Sonnet 5.5 is our second model in the Claude 5.5 family after Opus 5.5. It's a clear upgrade over Sonnet 5 and is smarter, more efficient and 30% faster. The per-token price is unchanged and because Sonnet 5.5 typically needs far fewer tokens to do the same work, it costs up to 30% less for most work. FIG ALance Martin's code-to-painting demo: each model writes code that repaints the same photograph. From left: the photograph, Claude Sonnet 5, Claude Sonnet 5.5 and Claude Opus 5.5. Credit

Hacker News:AI 热帖AI score 86/100

Prompting Claude Opus 5.5

This guide covers the prompting patterns specific to Claude Opus 5.5. For the model's capabilities and API changes, see What's new in Claude Opus 5.5. For techniques that apply across all current Claude models, see Prompting best practices. Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and tends to finish the same task with fewer tokens. Existing Claude Opus 5 prompts should perform well without changes, and the patterns in Prompting Claude Opus 5 remain

Every:最新文章(网页)AI score 50/100

Vibe Check: Sonnet 5.5 Finds Its Place in Claude’s Crowded Family

Anthropic's new model is a more capable partner than Sonnet 5, so long as you keep a hand on the wheel and know when to call in OpusSeptember 28, 2026Sonnet 5.5 is out today, and as a fellow middle child, I feel for it—it has something to prove. Subscribers onlyOnly available for paid subscribersSubscribe to read the rest of this article.Start free trial →Already have an account? Log in.Katie Parrott is a staff writer. She writes Working Overtime and contributes to Vibe Checks, Source Code, and

Claude Platform:开发者版本说明(RSS)AI score 71/100

Claude Platform release notes — September 28, 2026

The Claude Platform release notes list changes to the Claude API, the client SDKs, and the Claude Console, newest first. September 30, 2026 We announced the deprecation of the Claude Sonnet 4.5 model (claude-sonnet-4-5-20250929), with retirement on the Claude API scheduled for November 30, 2026. We recommend migrating to Claude Sonnet 5.5. Read more in Model deprecations. September 28, 2026 We've launched Claude Sonnet 5.5 (claude-sonnet-5-5). It's available on the Claude API, Claude in Amazon B

TechCrunch:AI(RSS)AI score 65/100

Anthropic’s CEO is about to have dinner with President Trump

Anthropic CEO Dario Amodei seems to be everywhere this weekend: He was lampooned on the season premiere of Saturday Night Live, and tonight, he’s set to have dinner with President Donald Trump at the White House.Axios first broke the news of Amodei’s dinner plans, which TechCrunch subsequently confirmed with a source familiar with his plans.This will be the first one-on-one meeting between the two men, who recently found themselves on opposite sides of the AI safety debate. Amodei released a pla

TechCrunch:AI(RSS)AI score 36/100

Anthropic’s Dario Amodei gets the ‘SNL’ treatment

In Brief Posted: 9:30 AM PDT · September 27, 2026 Image Credits:Saturday Night Live / NBC “Saturday Night Live” took on the AI industry’s recent warnings of doom last night, as cast member Jane Wickline offered her impression of Anthropic CEO Dario Amodei. Weekend Update host Michael Che — who sounded a little uncertain about how to pronounce Amodei’s last name — kicked the segment off by describing the CEO as having “stumbled through a press tour” where he seemingly agreed with a former employe

Fireworks AI(网页)AI score 62/100

Introducing FireRouter with Opus

Fireworks Nexus enables engineering teams to drop leading open models in the harnesses they already use and cut spend in half without sacrificing speed or quality. The solution includes FireRouter, the first cache-aware router on the market, which makes a big difference in speed and cost. Today, we’re introducing FireRouter with Opus, optimized for the Opus family and now available in both our CLI and, for the first time, as a standalone router model. Any Fireworks account can point to it as a s

NVIDIA Technical Blog:Agentic AI / Generative AIAI score 78/100

Add Runtime Controls to AI Agents with NVIDIA OpenShell

AI agents can be given a goal, write code, use tools, and keep working as new information becomes available. This opens the door to applications that investigate software failures, run experiments, and carry out business-critical actions and research over days or weeks. Useful agents need access to workspaces, compute resources, data, credentials, and external services. But broader access also creates more consequential failure modes, from changing production data or exposing confidential inform

Claude:Blog(网页)AI score 72/100

Giving companies more control over their AI agents, with NVIDIA

CategoryProductDateSeptember 28, 2026Reading time5minShareCopy linkhttps://claude.com/blog/giving-companies-more-control-over-their-ai-agents-with-nvidiaNVIDIA today announced the Open Agent Safety Platform, an open software platform and reference system design for strengthening AI security. Anthropic has collaborated with NVIDIA to bring additional layers of security and control to the agent stack. Claude Managed Agents, a suite of composable APIs for building and deploying production-grade age

Artificial Analysis 完整文章(网页)AI score 80/100

Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index

Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we've seenSee model page With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and moves to #2 on the Intelligence Index, behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens. However, it outputs a higher number of Output Tokens p

4 stories
The Decoder:AI News(RSS)AI score 55/100

Some Anthropic veterans are reportedly buying remote land in case "AI goes awry"

Some of Anthropic's longest-serving employees are making concrete contingency plans in case AI spirals out of control. Over the past few weeks, "some of its earliest employees" told an industry colleague they're considering buying land in remote parts of the US where they could relocate "if AI goes awry," the WSJ reports. The idea of fleeing AI isn't new in this crowd. Former employees recall that even at company dinners during Anthropic's early days, people discussed a Manhattan Project-style s

HuggingFace Daily Papers(社区热门论文)AI score 42/100

CompoWorld: Compositional Environment Scaling for General Agents

Authors:Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan View PDF HTML (experimental) Abstract:Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Composition

Simon Willison 博客AI score 62/100

Kākāpō Party

Tool Kākāpō Party — Experience an interactive pixel-art celebration featuring kākāpō (New Zealand's flightless parrots) jumping and dancing to music. Click, tap, or press Space to trigger confetti bursts, balloons, streamers, and festive effects while the birds party under a disco ball with choreographed dance moves that sync to an upbeat rhythm. I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the

9 stories
Hacker News:AI 热帖AI score 64/100

Drawgent: Coding agent on a live Excalidraw canvas

Rust 63.8% JavaScript 31.9% CSS 2.4% Shell 0.8% Makefile 0.6% HTML 0.2% Nix 0.2% Dockerfile 0.2% 50 1 1 Clone this repository Use permalink HTTPShttps://tangled.org/yanndegat.tngl.sh/drawgent https://tangled.org/did:plc:3zklierfkckthy6sodbm4af3 [email protected]:yanndegat.tngl.sh/drawgent [email protected]:did:plc:3zklierfkckthy6sodbm4af3 For self-hosted knots, clone URLs may differ based on your setup. Download tar.gz Download .zip README.md drawgent: your coding agent on a live Excalidraw canva

Hacker News:AI 热帖AI score 66/100

Show HN: A Claude Code skill to analyze your chess games

Claude Code skills that turn one of your chess games into a post-mortem you can actually read: plain-language explanations of your mistakes, checked against Stockfish, and a narrated video of the whole game. Edit: Damned this hit front page of HackerNews. You can go there for insightful conversations: https://news.ycombinator.com/item?id=49857528 I should have selected a game without a massive blunder on my side ahah. The video out-010.mp4 If the player above does not load (it needs a GitHub log

Hacker News:AI 热帖AI score 65/100

Understanding the Impact of LLM Watermarking on AI Agent Behavior

Recently, Anthropic announced that future Claude models would embed an invisible watermark in their output [1], [2], and subsequently disclosed that the watermark is based on Google DeepMind’s SynthID-Text [2], [3]. Text watermarking itself is not new, but its deployment now has regulatory relevance. Article 50(2) of the EU AI Act [4] requires providers of AI systems generating synthetic text to mark their outputs in a machine-readable format and make them detectable as artificially generated or

The Decoder:AI News(RSS)AI score 61/100

Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

A new study from Nvidia researchers tackles these costs not at the model level but at the harness, the control layer between the model and its environment used by systems like Codex, Claude Code, or OpenClaw. A research AI analyzes agent traces, proposes harness changes, and keeps only those that maintain performance while cutting costs. SoL-Pi saves 50 percent compared to Codex and 54.3 percent compared to Claude Code on EdgeBench. | Image: Nvidia The harness controls how an agent sees states,

TechCrunch:AI(RSS)AI score 73/100

Anthropic to pay Akamai $11.6 billion over seven years in cloud deal

Title: Anthropic to pay Akamai $11.6 billion over seven years in cloud deal URL Source: http://techcrunch.com/2026/09/25/anthropic-to-pay-akamai-11-6-billion-over-seven-years-in-cloud-deal Published Time: 2026-09-25T19:13:38+00:00 Markdown Content: Anthropic will spend $11.6 billion over seven years on Akamai’s cloud infrastructure, Akamai said Thursday. That’s more than six times the size of a $1.8 billion deal between the two companies that Bloomberg reported in May. The commitment isn’t ironc

The Decoder:AI News(RSS)AI score 76/100

Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says

Sep 25, 2026 A federal appeals court in Washington has upheld the Pentagon's decision to bar AI startup Anthropic from military contracts. The court ruled 2-1 that the Pentagon was right to classify Anthropic as a national security supply chain risk, CNBC reports. The case stems from Anthropic's refusal to allow its technology to be used for autonomous weapons and mass surveillance. Defense Secretary Pete Hegseth argued that Anthropic's safety restrictions could jeopardize military operations. A

Hacker News:AI 热帖AI score 68/100

Yes, Claude can do nine loops

In this guest post, physicist and science writer Matt von Hippel shares what happened when he issued a challenge to AI companies regarding a problem in his former subfield of theoretical physics.It’s not often that you issue a challenge, only to see it beaten a month later. But we’re living in unusual times.Let me introduce myself: I’m Matt von Hippel. I used to be a theoretical physicist; these days I’m a science writer. Throughout, I’ve been a blogger, writing weekly at 4gravitons.com about ph

Hacker News:AI 热帖AI score 67/100

Jevmem - automatic project memory for Claude Code, built on Jev

Say it once. jevmem writes down what you decide in Claude Code and brings it back next session. In Anthropic's Claude plugin directory · On the MCP Registry · Open source, MIT jevmem-0.6-readme.mp4 Why Each Claude Code session starts with a fresh context, so what you decided last week lives in last week's chat. A CLAUDE.md file helps if you keep it up to date. jevmem keeps a file like it up to date for you, as you work. How it works You decide something in a chat: "We use Postgres." jevmem asks

5 stories
TechCrunch:AI(RSS)AI score 76/100

Anthropic’s founders seek voting control ahead of IPO

In Brief Posted: 8:40 AM PDT · September 25, 2026 Image Credits:Samyukta Lakshmi/Bloomberg / Getty Images Anthropic’s founders want to stay in charge after the company goes public. According to The Information, the company is asking shareholders to approve a structure “in the coming days” that would give CEO Dario Amodei and his six co-founders special shares carrying a combined 50.1% of the vote on most corporate matters, as long as at least three of them keep a minimum stake. This isn’t a new

Hacker News:AI 热帖AI score 79/100

U.S. appeals court upholds designation of Anthropic as supply chain risk

Title: U.S. appeals court upholds Pentagon designation of Anthropic as supply chain risk URL Source: http://www.cnbc.com/2026/09/25/pentagon-anthropic-ai-risk-appeals-court.html Published Time: 2026-09-25T15:25:01+0000 Markdown Content: watch now A federal appeals court panel in Washington, D.C., on Friday upheld the Pentagon’s blacklisting of Anthropic, dealing a blow to the artificial intelligence company in its months-long battle with the Trump administration. The 2-1 decision rejected Anthro

4 stories
3 stories
4 stories
1 story
1 story
Showing the latest 86 of 658 stories.