摘要
Automated Alignment Researcher(AAR)是一个固定的自动研究 harness:它让多个 LLM researcher agents 搜索文献、提出 post-training 方法、提交代码和模型,并依据隔离 benchmark 持续筛选候选。(来源:Automated Researchers Can Reliably Mitigate Alignment Failures)
核心内容
基本信息
- 实体类别:自动研究 harness / 实验系统
- 关键名称或版本:Claude Opus 4.8 AAR;Claude Sonnet 5 production-grade pilot
- 时间范围:2026-08-28 首次公开论文
重要事实
- 运行时由 librarian survey、五个并行 researcher、fresh sessions、finding forum/leaderboard、代码 monitor、训练 worker 与隔离 evaluator 组成。(来源:Automated Researchers Can Reliably Mitigate Alignment Failures,图 2)
- 候选方法必须提交模型权重并通过能力门;研究侧不能读取 hidden held-out data,也不能从 AAR 或更强模型直接蒸馏。(来源:Automated Researchers Can Reliably Mitigate Alignment Failures,§2-3)
- AAR 本身不通过 RL 更新,也不修改其 monitor/evaluator,因此不是严格 Agentic RL 或“改进 improver”层 RSI。(综合:Automated Researchers Can Reliably Mitigate Alignment Failures)
关系与影响
AAR 把 Harness Optimization 的可见搜索过程用于模型 post-training,而非直接修改 harness;其共享结果、隔离验证和反作弊机制可作为自动研究系统的控制面参考。
分歧与变化
作者将结果解释为“可测 alignment failure 上自动 post-training 的早期可行性”,但该范围不等于开放 alignment research,也不证明通用 Recursive Self-Improvement。(来源:Automated Researchers Can Reliably Mitigate Alignment Failures,§8)
相关页面
- 概念:Harness Optimization、Agentic RL、Recursive Self-Improvement、Self-Improving Harness
- 实体:无
- 主题:LLM Agent Self-Improvement
来源
待核实问题
- 公开仓库缺少许可证,且完整 48 小时研究 loop 与 held-out 环境尚待独立复现。