GPTProto

Qwen / Alibaba News

Latest reporting, research and product updates filed under Qwen / Alibaba.

76 picksNewest firstLatest article Sep 23, 2026, 5:11 PM

Latest stories

2 stories
2 stories
2 stories
9 stories
HuggingFace Daily Papers(社区热门论文)AI score 39/100

Rubric Rewards from Item Response Theory

View PDF HTML (experimental) Abstract:Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguish

HuggingFace Daily Papers(社区热门论文)AI score 38/100

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Authors:Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing View PDF HTML (experimental) Abstract:Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to

HuggingFace Daily Papers(社区热门论文)AI score 39/100

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

View PDF HTML (experimental) Abstract:Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as re

HuggingFace Daily Papers(社区热门论文)AI score 56/100

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

View PDF HTML (experimental) Abstract:Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to

HuggingFace Daily Papers(社区热门论文)AI score 62/100

Language Models Are "Insecure" Reporters

View PDF HTML (experimental) Abstract:As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine a

HuggingFace Daily Papers(社区热门论文)AI score 32/100

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

View PDF HTML (experimental) Abstract:Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one

HuggingFace Daily Papers(社区热门论文)AI score 40/100

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

View PDF HTML (experimental) Abstract:Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens.

HuggingFace Daily Papers(社区热门论文)AI score 42/100

LLMs are General Asynchronous Agents

View PDF HTML (experimental) Abstract:Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling fo

HuggingFace Daily Papers(社区热门论文)AI score 44/100

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

View PDF HTML (experimental) Abstract:Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating fine

1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
Showing the latest 36 of 76 stories.