◇ 今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

人读卡片 · 五问

① 这是什么

今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。

② 值不值得用

进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。

③ 怎么开始

先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。

展开:④ 关键证据 / ⑤ 注意事项与坑

④ 关键证据

  • 论文 ID 2608.14905
  • 作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa

⑤ 注意事项与坑

摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。

一句话使用前提

适合:追踪今日 AI 论文前沿的读者。

不适合:需要完整复现实验细节的场景(请读原文)。

状态:机检通过(结构与溯源校验);功能未逐项实测(未校验)。

最小验证(3 步)

  1. 读摘要 — How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
  2. 核作者/日期 — Yanlin Fei, Nazhou Liu, Xinmiao Yu · 2026-08-14
  3. 跑 CLI — npx aiskillready search 14905
快速开始

可以这样用

npx aiskillready search 14905   # 本站 CLI:真实可跑(aiskillready@0.2.0)

说明:`npx aiskillready` 是本站已发布 CLI(npm aiskillready@0.2.0),命令真实可跑;第三方命令按官方文档给出、未逐一实测,本平台不编造结论。

打开来源 ↗
信任层:机检 / 成本 / 置信度 / schema(展开查看)

条目 ID:ASR-PAPER-20260818-2608-14905 | 类型:论文 | schema v0.5

✅ 机检 通过 💰 升维成本 $0.0000 🧭 置信度 中 📌 schema 0.5 🛂 审核 agent reviewed 🧪 功能未验证(Top20 实测计划)

相关推荐(应该知道的)

相关推荐(3 条)· 展开更多
反馈

这条不对 / 我执行报错了?带 entry_id 一键提 Issue,回流后修订版本。

🐞 纠错 / 报错 → 提 Issue
AI 区块 · 机器可读(同一事实源)
AI 区块 · 机器可读 JSON(点击展开)
{
  "schema_version": "0.5",
  "entry_id": "ASR-PAPER-20260818-2608-14905",
  "record_type": "paper",
  "title": "今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
  "observed_at": "2026-08-18",
  "source": {
    "name": "huggingface.co/api/daily_papers",
    "url": "https://huggingface.co/api/daily_papers?limit=10",
    "snapshot_hash": "652c1f13fb67dd93f7d60fecc11565d77173419a1e4693958922f4b73738fb42",
    "fetched_at": "2026-08-18T10:51:18+00:00"
  },
  "confidence": "medium",
  "human_stream": {
    "summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowl",
    "key_points": [
      "论文 ID 2608.14905",
      "作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa",
      "发布 2026-08-14T00:00:00.000Z"
    ],
    "note": "来源为 HuggingFace daily papers 公开 API;完整阅读请走原文链接。",
    "what_it_is": "今日论文《今日论文 · How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks》(Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa),来自 HF Daily Papers 公开源。",
    "worth_it": "进入 HF 每日精选说明有社区关注度;是否与你的问题相关需要读原文判断,本条目不评价研究质量。",
    "how_to_start": "先读摘要,需要再读全文(https://huggingface.co/papers/2608.14905);对照论文 ID(2608.14905)可找实现与讨论。",
    "evidence": [
      "论文 ID 2608.14905",
      "作者 Yanlin Fei、Nazhou Liu、Xinmiao Yu、Shaolong Chen、Lei Li、Rahul Thapa"
    ],
    "cautions": "摘要来自作者原文,未做同行评议级核验;引用请以正式发表版本为准。"
  },
  "ai_stream": {
    "structured": {
      "paper_id": "2608.14905",
      "title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
      "authors": [
        "Yanlin Fei",
        "Nazhou Liu",
        "Xinmiao Yu",
        "Shaolong Chen",
        "Lei Li",
        "Rahul Thapa",
        "Madalina Ciobanu",
        "Qingqing Mao",
        "Ritankar Das"
      ],
      "published_at": "2026-08-14T00:00:00.000Z",
      "url": "https://huggingface.co/papers/2608.14905",
      "source_label": "hf-daily-papers"
    },
    "raw": [
      {
        "paper_id": "2608.14905",
        "title": "How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks",
        "authors": [
          "Yanlin Fei",
          "Nazhou Liu",
          "Xinmiao Yu",
          "Shaolong Chen",
          "Lei Li",
          "Rahul Thapa",
          "Madalina Ciobanu",
          "Qingqing Mao",
          "Ritankar Das"
        ],
        "summary": "AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 10",
        "published_at": "2026-08-14T00:00:00.000Z",
        "submitted_on_daily_at": null,
        "url": "https://huggingface.co/papers/2608.14905"
      }
    ]
  },
  "token_cost": {
    "total": 0.0,
    "currency": "USD",
    "breakdown": {
      "crawl": 0.0,
      "clean": 0.0,
      "elevate": 0.0,
      "verify": 0.0
    }
  },
  "machine_verified": true,
  "review": {
    "status": "agent_reviewed",
    "reviewer": "pipeline-validate",
    "reviewed_at": "2026-08-18T10:54:20+00:00",
    "comments": "机检通过(schema 0 error + 溯源一致 + 成本达标)"
  },
  "provenance": {
    "extracted_by": "codex",
    "extracted_at": "2026-08-18T10:53:05+00:00",
    "pipeline": "collect_hf_papers.py + build_entries.py v0.2(规则管线)",
    "access_urls": [
      "https://huggingface.co/api/daily_papers?limit=10"
    ],
    "card_generated_at": "2026-08-18T10:53:05+00:00",
    "card_pipeline": "enrich_human_stream.py v0.1(规则模板,零 Token)"
  }
}