◇ 今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
① 这是什么
今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。
② 值不值得用
进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。
③ 怎么开始
先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。
展开:④ 关键证据 / ⑤ 注意事项与坑
④ 关键证据
- 论文 ID 2608.14905
- 作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa
⑤ 注意事项与坑
摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。
适合:追踪今日 AI 论文前沿的读者。
不适合:需要完整复现实验细节的场景(请读原文)。
状态:机检通过(结构与溯源校验);功能未逐项实测(未校验)。
最小验证(3 步)
- 读摘要 — How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
- 核作者/日期 — Yanlin Fei, Nazhou Liu, Xinmiao Yu · 2026-08-14
- 跑 CLI — npx aiskillready search 14905
可以这样用
npx aiskillready search 14905 # 本站 CLI:真实可跑(aiskillready@0.2.0)
说明:`npx aiskillready` 是本站已发布 CLI(npm aiskillready@0.2.0),命令真实可跑;第三方命令按官方文档给出、未逐一实测,本平台不编造结论。
信任层:机检 / 成本 / 置信度 / schema(展开查看)
相关推荐(应该知道的)
相关推荐(3 条)· 展开更多
◇ 今日论文 · Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Mani
相似◇ 今日论文 · MOSS-VL Technical Report
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay
相似◇ 今日论文 · WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, y
这条不对 / 我执行报错了?带 entry_id 一键提 Issue,回流后修订版本。
🐞 纠错 / 报错 → 提 IssueAI 区块 · 机器可读 JSON(点击展开)
{
"schema_version": "0.5",
"entry_id": "ASR-PAPER-20260818-2608-14905",
"record_type": "paper",
"title": "今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"observed_at": "2026-08-18",
"source": {
"name": "huggingface.co/api/daily_papers",
"url": "https://huggingface.co/api/daily_papers?limit=10",
"snapshot_hash": "652c1f13fb67dd93f7d60fecc11565d77173419a1e4693958922f4b73738fb42",
"fetched_at": "2026-08-18T10:51:18+00:00"
},
"confidence": "medium",
"human_stream": {
"summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowl",
"key_points": [
"论文 ID 2608.14905",
"作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa",
"发布 2026-08-14T00:00:00.000Z"
],
"note": "来源为 HuggingFace daily papers 公开 API;完整阅读请走原文链接。",
"what_it_is": "今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。",
"worth_it": "进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。",
"how_to_start": "先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。",
"evidence": [
"论文 ID 2608.14905",
"作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa"
],
"cautions": "摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。"
},
"ai_stream": {
"structured": {
"paper_id": "2608.14905",
"title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"authors": [
"Yanlin Fei",
"Nazhou Liu",
"Xinmiao Yu",
"Shaolong Chen",
"Lei Li",
"Rahul Thapa",
"Madalina Ciobanu",
"Qingqing Mao",
"Ritankar Das"
],
"published_at": "2026-08-14T00:00:00.000Z",
"url": "https://huggingface.co/papers/2608.14905",
"source_label": "hf-daily-papers"
},
"raw": [
{
"paper_id": "2608.14905",
"title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"authors": [
"Yanlin Fei",
"Nazhou Liu",
"Xinmiao Yu",
"Shaolong Chen",
"Lei Li",
"Rahul Thapa",
"Madalina Ciobanu",
"Qingqing Mao",
"Ritankar Das"
],
"summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 10",
"published_at": "2026-08-14T00:00:00.000Z",
"submitted_on_daily_at": null,
"url": "https://huggingface.co/papers/2608.14905"
}
]
},
"token_cost": {
"total": 0.0,
"currency": "USD",
"breakdown": {
"crawl": 0.0,
"clean": 0.0,
"elevate": 0.0,
"verify": 0.0
}
},
"machine_verified": true,
"review": {
"status": "agent_reviewed",
"reviewer": "pipeline-validate",
"reviewed_at": "2026-08-18T10:54:20+00:00",
"comments": "机检通过(schema 0 error + 溯源一致 + 成本达标)"
},
"provenance": {
"extracted_by": "codex",
"extracted_at": "2026-08-18T10:53:05+00:00",
"pipeline": "collect_hf_papers.py + build_entries.py v0.2(规则管线)",
"access_urls": [
"https://huggingface.co/api/daily_papers?limit=10"
],
"card_generated_at": "2026-08-18T10:53:05+00:00",
"card_pipeline": "enrich_human_stream.py v0.1(规则模板,零 Token)"
}
}