◇ Paper · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
| Official / source | https://huggingface.co/papers/2608.14905 |
|---|
Chinese editor copy (original, collapsed)
① 这是什么
今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。
② 值不值得用
进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。
③ 怎么开始
先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。
④ 关键证据
- 论文 ID 2608.14905
- 作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa
⑤ 注意事项与坑
摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。
Good for: readers tracking today's AI research frontier.
Not for: reproducing full experimental details (read the paper).
Status: machine-checked (structure & source validation); functionality not individually tested.
Minimal verification (3 steps)
- Read abstract — How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
- Check authors/date — Yanlin Fei, Nazhou Liu, Xinmiao Yu · 2026-08-14
- Run CLI — npx aiskillready search 14905
Use it like this
npx aiskillready search 14905 # official CLI, verified runnable (aiskillready@0.2.0)
Note: `npx aiskillready` is this site's published CLI (npm aiskillready@0.2.0) — verified runnable; third-party commands are given per official docs and not individually tested. We don't fabricate conclusions.
Trust layer: machine check / cost / confidence / schema (expand)
Related (what you should know)
Related (3) · expand
◇ Paper · Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift. by Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi. published 2026-08-15.
Similar◇ Paper · MOSS-VL Technical Report
MOSS-VL Technical Report. by Pengyu Wang, Chenkun Tan, Shaojun Zhou, Qirui Zhou, Yanxin Chen, Xingyang He, Huazheng Zeng, Jijun Cheng, Chenghao Wang,. published 2026-08-15.
Similar◇ Paper · WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations. by Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge. published 2026-08-16.
Something wrong / command failed? File an issue with the entry ID; it feeds back into the next revision.
🐞 Report an issue →AI block · machine-readable JSON (click to expand)
{
"schema_version": "0.5",
"entry_id": "ASR-PAPER-20260818-2608-14905",
"record_type": "paper",
"title": "今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"observed_at": "2026-08-18",
"source": {
"name": "huggingface.co/api/daily_papers",
"url": "https://huggingface.co/api/daily_papers?limit=10",
"snapshot_hash": "652c1f13fb67dd93f7d60fecc11565d77173419a1e4693958922f4b73738fb42",
"fetched_at": "2026-08-18T10:51:18+00:00"
},
"confidence": "medium",
"human_stream": {
"summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowl",
"key_points": [
"论文 ID 2608.14905",
"作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa",
"发布 2026-08-14T00:00:00.000Z"
],
"note": "来源为 HuggingFace daily papers 公开 API;完整阅读请走原文链接。",
"what_it_is": "今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。",
"worth_it": "进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。",
"how_to_start": "先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。",
"evidence": [
"论文 ID 2608.14905",
"作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa"
],
"cautions": "摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。"
},
"ai_stream": {
"structured": {
"paper_id": "2608.14905",
"title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"authors": [
"Yanlin Fei",
"Nazhou Liu",
"Xinmiao Yu",
"Shaolong Chen",
"Lei Li",
"Rahul Thapa",
"Madalina Ciobanu",
"Qingqing Mao",
"Ritankar Das"
],
"published_at": "2026-08-14T00:00:00.000Z",
"url": "https://huggingface.co/papers/2608.14905",
"source_label": "hf-daily-papers"
},
"raw": [
{
"paper_id": "2608.14905",
"title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
"authors": [
"Yanlin Fei",
"Nazhou Liu",
"Xinmiao Yu",
"Shaolong Chen",
"Lei Li",
"Rahul Thapa",
"Madalina Ciobanu",
"Qingqing Mao",
"Ritankar Das"
],
"summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 10",
"published_at": "2026-08-14T00:00:00.000Z",
"submitted_on_daily_at": null,
"url": "https://huggingface.co/papers/2608.14905"
}
]
},
"token_cost": {
"total": 0.0,
"currency": "USD",
"breakdown": {
"crawl": 0.0,
"clean": 0.0,
"elevate": 0.0,
"verify": 0.0
}
},
"machine_verified": true,
"review": {
"status": "agent_reviewed",
"reviewer": "pipeline-validate",
"reviewed_at": "2026-08-18T10:54:20+00:00",
"comments": "机检通过(schema 0 error + 溯源一致 + 成本达标)"
},
"provenance": {
"extracted_by": "codex",
"extracted_at": "2026-08-18T10:53:05+00:00",
"pipeline": "collect_hf_papers.py + build_entries.py v0.2(规则管线)",
"access_urls": [
"https://huggingface.co/api/daily_papers?limit=10"
],
"card_generated_at": "2026-08-18T10:53:05+00:00",
"card_pipeline": "enrich_human_stream.py v0.1(规则模板,零 Token)"
}
}