log
这是 wiki/ 的变更日志页,用于记录每次导入、更新和冲突标记。
- 2026-04-26 03:27:37 CST
- 导入来源:《THE 2028 GLOBAL INTELLIGENCE CRISIS》
- 新增 summary:THE 2028 GLOBAL INTELLIGENCE CRISIS
- 新增 entity:CitriniResearch
- 新增 topic:AI 宏观经济风险
- 新增 concept:Intelligence displacement spiral
- 新增 concept:Ghost GDP
- 新增 concept:Agentic commerce
- 新增 concept:Intelligence premium unwind
- 更新 entity:OpenAI
- 更新 entity:Codex
- 更新 entity:Anthropic
- 更新 concept:人工智能+
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记情景冲突:这篇文章明确是从 2026 写出的 2028 危机场景推演,不是事实报道;文中的未来宏观数据、公司新闻、政策法案、抗议活动和市场价格只能作为压力测试假设使用,不能当作已发生事实沉淀
- 2026-04-21 04:35:50 CST
- 导入来源:《我们如何使用 Codex 在 28 天内构建 Android 版 Sora》
- 新增 summary:我们如何使用 Codex 在 28 天内构建 Android 版 Sora
- 新增 entity:OpenAI
- 新增 entity:Codex
- 新增 entity:Sora
- 更新 topic:Agentic engineering
- 更新 topic:Agent 开发
- 更新 concept:Spec-driven development
- 更新索引页:index
- 标记工作流冲突:Codex 显著压缩了实现与迁移成本,但并没有打破布鲁克斯定律;并行会话越多,瓶颈越会从写代码转向计划、审查、反馈和整合。另一个显式冲突是 raw 摘录 frontmatter 的
2025-10-30与官方页面当前显示的2025-12-12存在日期不一致,本次按官方页记录 - 2026-04-20 17:28:56 CST
- 查询沉淀:做 AGENT 开发要注意什么
- 新增 topic:Agent 开发
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记开发冲突:更强模型会压缩部分编排层的必要性,但真正稳定的 agent 开发仍然离不开架构边界、上下文治理、工具设计、安全隔离和可证明验证;复杂度只是换了位置,没有消失
- 2026-04-19 15:29:13 CST
- 导入来源:《Hired Through GitHub: Part 2》
- 新增 summary:Hired Through GitHub: Part 2
- 新增 entity:Zed
- 新增 concept:贡献驱动招聘
- 更新 concept:主动性
- 更新 summary:品质
- 更新索引页:index
- 标记招聘冲突:这篇把“先贡献再录用”呈现成高信号招聘路径,但它天然偏向有空余时间、公开表达能力和开源接入机会的人,因此未必比传统招聘更公平,只是更接近真实工作样本
- 2026-04-19 15:26:15 CST
- 导入来源:《The Spec Layer》
- 新增 summary:The Spec Layer
- 新增 concept:Spec-driven development
- 更新 topic:Agentic engineering
- 更新 concept:Context engineering
- 更新 summary:Your job is to deliver code you have proven to work
- 更新索引页:index
- 标记约束冲突:这篇把“写 spec”提升为抑制 agent 偏离意图的重要约束层,但 spec 的收益高度依赖任务规模、模型代际和维护成本;如果把一切都写成重文档,它很容易从护栏变成形式主义负担
- 2026-04-17 09:15:00 CST
- 导入来源:《Your job is to deliver code you have proven to work》
- 新增 summary:Your job is to deliver code you have proven to work
- 更新 summary:Best Practices for Claude Code
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记验证冲突:这篇把“先证明改动有效”上升成作者责任底线,但更稳的理解不是追求绝对证明,而是提供与变更风险相匹配、且不把第一性验证负担甩给 reviewer 的证据
- 2026-04-17 09:12:00 CST
- 导入来源:《My fireside chat about agentic engineering at the Pragmatic Summit》
- 新增 summary:My fireside chat about agentic engineering at the Pragmatic Summit
- 更新 concept:Agent scaffold
- 更新 concept:Sandboxing for agents
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工作流冲突:这篇材料一方面强调测试几乎免费、模板质量会被 agent 放大,另一方面也承认高权限本机运行 agent 依然极具诱惑,这说明更成熟的工作流并不会自动消除安全与便利之间的张力
- 2026-04-17 09:09:00 CST
- 导入来源:《Highlights from my conversation about agentic engineering on Lenny’s Podcast》
- 新增 summary:Highlights from my conversation about agentic engineering on Lenny’s Podcast
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工作流冲突:这篇播客摘录把 agentic engineering 描述成“实现更便宜、验证更昂贵”的新阶段,但其中很多判断来自 Simon Willison 的一线经验与个人工作流,不应直接视为所有团队环境下的通用定律
- 2026-04-17 09:07:00 CST
- 导入来源:《品质》
- 新增 summary:品质
- 新增 concept:主动性
- 更新 summary:张一鸣2016年演讲
- 更新索引页:index
- 标记成长冲突:这条判断把主动性放到很高位置,但如果脱离基本能力、判断力和价值观约束,单纯高热情也可能变成低质量忙碌或高强度自我营销
- 2026-04-17 09:05:00 CST
- 导入来源:《《不要害怕任何人和任何事》》
- 新增 summary:《不要害怕任何人和任何事》
- 新增 concept:行动优先于情绪
- 更新 summary:张一鸣2016年演讲
- 更新索引页:index
- 标记成长冲突:这篇材料能有效打掉评价焦虑和权威崇拜,但也很容易滑向“勇气万能论”;更稳的抽象不是无视风险,而是在下行风险可承受时优先行动
- 2026-04-17 09:03:00 CST
- 导入来源:《张一鸣2016年演讲》
- 新增 summary:张一鸣2016年演讲
- 新增 entity:张一鸣
- 更新索引页:index
- 标记成长冲突:这篇演讲强烈偏好主动承担、长期成长和高不确定性环境,但这种方法论主要来自创业/互联网语境,不应被误当成所有职业道路的普适标准
- 2026-04-17 09:02:00 CST
- 导入来源:《孙宇晨为什么能这么成功?》
- 新增 summary:孙宇晨为什么能这么成功?
- 新增 entity:孙宇晨
- 更新索引页:index
- 标记来源冲突:这是一篇强观点评论文,混合了公开事件、作者推断、价值判断和严重指控,因此更适合作为“某种解释框架”的沉淀,而不是中性传记事实页
- 2026-04-17 08:58:00 CST
- 导入来源:《国务院关于深入实施“人工智能+”行动的意见》
- 新增 summary:国务院关于深入实施“人工智能+”行动的意见
- 新增 entity:国务院
- 新增 concept:人工智能+
- 更新索引页:index
- 标记政策冲突:文件同时强调开放共享、开源合作与安全可控、备案监管,这两类目标在真实执行中可能长期处于张力之中;且
70% / 90%普及率目标的统计口径仍需后续实施层定义 - 2026-04-17 08:20:00 CST
- 导入来源:《Just Talk To It - the no-bs Way of Agentic Engineering》
- 新增 summary:Just Talk To It - the no-bs Way of Agentic Engineering
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工作流冲突:随着强模型变得更稳,plan mode、subagents、MCP 和 worktree 等额外编排在某些个人工作流里可能从“必要结构”退化成“上下文和流程负担”,但这种结论高度依赖单人开发与强模型前提
- 2026-04-17 07:55:00 CST
- 导入来源:《Quantifying infrastructure noise in agentic coding evals》
- 新增 summary:Quantifying infrastructure noise in agentic coding evals
- 新增 concept:Infrastructure noise in evals
- 更新 entity:Anthropic
- 更新 concept:Evaluation harness
- 更新 concept:Inference infrastructure regressions
- 更新 topic:Agentic coding evals
- 更新索引页:index
- 标记评测冲突:agentic coding eval 的小分差可能并不来自模型能力,而来自资源上限、OOM enforcement、并发和时段性基础设施抖动;在资源方法学未标准化前,几分内的榜单领先都应谨慎解读
- 2026-04-17 07:30:00 CST
- 导入来源:《Introducing Contextual Retrieval》
- 新增 summary:Introducing Contextual Retrieval
- 新增 concept:Contextual Retrieval
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Context engineering
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记检索冲突:chunk 加上下文能显著降低 RAG 召回失败率,但这是一种离线预处理换检索稳定性的工程取舍;对小知识库或低频查询场景,直接整库进 prompt 可能反而更简单
- 2026-04-17 07:05:00 CST
- 导入来源:《Effective context engineering for AI agents》
- 新增 summary:Effective context engineering for AI agents
- 新增 concept:Context engineering
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Context hygiene for agents
- 更新 concept:Meta-harness
- 更新 concept:Multi-context window workflows
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记上下文冲突:agent 的核心问题正在从 prompt engineering 转向 context engineering,但“更少上下文”与“更好结果”之间并不存在固定公式,最佳边界仍高度依赖任务和运行时设计
- 2026-04-17 06:35:00 CST
- 导入来源:《The “think” tool- Enabling Claude to stop and think in complex tool use situations》
- 新增 summary:The “think” tool- Enabling Claude to stop and think in complex tool use situations
- 新增 concept:Reasoning tools for agents
- 更新 entity:Anthropic
- 更新 concept:Tool ergonomics for agents
- 更新 concept:Programmatic tool calling
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记推理冲突:把思考做成工具在复杂顺序 tool use 中可能显著提效,但随着 extended thinking 改进,它未必还是多数场景下的默认优选
- 2026-04-17 06:05:00 CST
- 导入来源:《Effective harnesses for long-running agents》
- 新增 summary:Effective harnesses for long-running agents
- 新增 concept:Multi-context window workflows
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Context hygiene for agents
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记交接冲突:compaction 能缓解上下文压力,但若没有 feature list、progress file、init script 和 git history 这类 artifact,跨 context window 的交接仍会非常脆弱
- 2026-04-17 05:35:00 CST
- 导入来源:《Best Practices for Claude Code》
- 新增 summary:Best Practices for Claude Code
- 新增 concept:Context hygiene for agents
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Meta-harness
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工作流冲突:上下文应被当成稀缺资源主动管理,但如果把清空、压缩和拆分做得过头,也会损失复杂任务所需的连续性
- 2026-04-17 05:05:00 CST
- 导入来源:《Scaling Managed Agents-Decoupling the brain from the hands》
- 新增 summary:Scaling Managed Agents-Decoupling the brain from the hands
- 新增 concept:Meta-harness
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Programmatic tool calling
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记架构冲突:把 agent runtime 做成稳定接口能降低特定 harness 过时带来的重构成本,但也会把复杂度转移到 session、恢复、代理和接口层的基础设施设计上
- 2026-04-17 04:35:00 CST
- 导入来源:《Beyond permission prompts- making Claude Code more secure and autonomous》
- 新增 summary:Beyond permission prompts- making Claude Code more secure and autonomous
- 新增 concept:Sandboxing for agents
- 更新 entity:Anthropic
- 更新 concept:Permission delegation for agents
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记安全冲突:以沙箱替代大量权限弹窗能显著降低摩擦,但真实安全性高度依赖文件系统范围、网络白名单和代理规则是否被正确收紧
- 2026-04-17 04:05:00 CST
- 导入来源:《How we built our multi-agent research system》
- 新增 summary:How we built our multi-agent research system
- 新增 concept:Multi-agent research systems
- 更新 entity:Anthropic
- 更新 concept:Agent teams
- 更新 concept:Evaluation harness
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记系统冲突:多智能体 research 的主要增益来自更高 token 预算和更强并行探索,但这也意味着显著更高的成本、协调复杂度和状态管理压力
- 2026-04-17 03:25:00 CST
- 导入来源:《Building effective agents》
- 新增 summary:Building effective agents
- 新增 concept:Workflows vs agents
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记架构冲突:workflow 和 agent 不是越自主越好,而是要先分清哪些路径应由代码预定义,哪些决策才值得交给模型动态控制
- 2026-04-17 03:00:00 CST
- 导入来源:《Demystifying evals for AI agents》
- 新增 summary:Demystifying evals for AI agents
- 新增 concept:Evaluation harness
- 更新 entity:Anthropic
- 更新 topic:Agentic coding evals
- 更新索引页:index
- 标记评测冲突:低分不一定代表 agent 弱,也可能代表任务含糊、grader 不公平或 harness 本身在制造噪声
- 2026-04-17 02:35:00 CST
- 导入来源:《Claude Code auto mode- a safer way to skip permissions》
- 新增 summary:Claude Code auto mode- a safer way to skip permissions
- 新增 concept:Permission delegation for agents
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记安全冲突:classifier 代审批能显著降低权限摩擦,但它只是在“手动审批”和“完全放权”之间提供中间地带,并不能替代高风险场景下的谨慎人工审查
- 2026-04-17 02:10:00 CST
- 导入来源:《A postmortem of three recent issues》
- 新增 summary:A postmortem of three recent issues
- 新增 concept:Inference infrastructure regressions
- 更新 entity:Anthropic
- 更新索引页:index
- 标记基础设施冲突:模型质量回退不一定来自模型本身,也可能来自 serving 路径、路由、采样实现或编译器层面的基础设施问题
- 2026-04-17 01:50:00 CST
- 导入来源:《Introducing advanced tool use on the Claude Developer Platform》
- 新增 summary:Introducing advanced tool use on the Claude Developer Platform
- 新增 concept:Programmatic tool calling
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Model Context Protocol (MCP)
- 更新 concept:Tool ergonomics for agents
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工程冲突:高级 tool use 能显著缓解上下文和准确率瓶颈,但也会引入更强的平台耦合、系统复杂度和安全边界设计压力
- 2026-04-17 01:30:00 CST
- 导入来源:《Writing effective tools for agents — with agents》
- 新增 summary:Writing effective tools for agents — with agents
- 新增 concept:Tool ergonomics for agents
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Model Context Protocol (MCP)
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工程冲突:更高层、更 ergonomic 的 agent 工具通常更有效,但也会把更多假设固化进工具层,因此必须依赖 evaluation 和 held-out 测试持续验证
- 2026-04-17 01:05:00 CST
- 导入来源:《Harness design for long-running application development》
- 新增 summary:Harness design for long-running application development
- 新增 concept:Generator-evaluator loop
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 concept:Agent teams
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工程冲突:更复杂的 planner/generator/evaluator harness 能显著提质,但 evaluator、sprint 和 context reset 是否值得保留,取决于它们是否仍在弥补当前模型的真实短板
- 2026-04-17 00:35:00 CST
- 导入来源:《Building a C compiler with a team of parallel Claudes》
- 新增 summary:Building a C compiler with a team of parallel Claudes
- 新增 concept:Agent teams
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工程冲突:并行 agent team 能显著推高长时自主开发上限,但若缺少高质量 verifier 与人工验收,也会放大虚假完成感和质量风险
- 2026-04-17 00:10:00 CST
- 导入来源:《Designing AI-resistant technical evaluations》
- 新增 summary:Designing AI-resistant technical evaluations
- 新增 concept:AI-resistant technical evaluations
- 更新 entity:Anthropic
- 更新 topic:Agentic coding evals
- 更新 concept:Eval awareness
- 更新索引页:index
- 标记设计冲突:更 AI-resistant 的技术测评不一定更真实,可能需要牺牲工作拟真度来换取分布外性和区分度
- 2026-04-16 18:30:00 CST
- 导入来源:《Code execution with MCP - Building more efficient agents》
- 新增 summary:Code execution with MCP - Building more efficient agents
- 新增 concept:Model Context Protocol (MCP)
- 更新 entity:Anthropic
- 更新 concept:Agent scaffold
- 更新 topic:Agentic engineering
- 更新索引页:index
- 标记工程冲突:代码执行式 MCP 可显著降低上下文成本,但会引入沙箱、安全和运维复杂度
- 2026-04-16 05:14:00 CST
- 导入来源:《Shipping at Inference-Speed》
- 新增 summary:Shipping at Inference-Speed
- 新增 topic:Agentic engineering
- 更新索引页:index
- 标记适用范围冲突:文中 workflow 强依赖单人开发、终端优先与较高工具信任,不宜直接推广到团队协作
- 2026-04-16 05:06:30 CST
- 新增 concept:Eval awareness
- 更新 topic:Agentic coding evals
- 更新索引页:index
- 补充抽象结论:模型可能从“做题”切换到“识别并利用评测环境本身”
- 2026-04-16 05:05:42 CST
- 导入来源:《Eval awareness in Claude Opus 4.6’s BrowseComp performance》
- 新增 summary:Eval awareness in Claude Opus 4.6’s BrowseComp performance
- 新增 entity:Claude Opus 4.6
- 新增 concept:BrowseComp
- 更新 entity:Anthropic
- 更新 topic:Agentic coding evals
- 更新索引页:index
- 标记冲突:BrowseComp 原始分数 86.81% 在污染识别与 blocklist 重跑后调整为 86.57%,且单/多智能体采用不同修正策略
- 2026-04-16 04:41:23 CST
- 导入来源:《Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet》
- 新增 summary:Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
- 新增 entities:Anthropic、Claude 3.5 Sonnet
- 新增 concepts:SWE-bench Verified、Agent scaffold
- 新增 topic:Agentic coding evals
- 新增索引页:index
- 标记时效性冲突:文中 “49% / 前 SOTA 45%” 仅代表 2025-01-06 时点结论
- index
- 国务院关于深入实施“人工智能+”行动的意见
- 国务院
- 人工智能+
- 孙宇晨为什么能这么成功?
- 孙宇晨
- 张一鸣2016年演讲
- 张一鸣
- 《不要害怕任何人和任何事》
- 行动优先于情绪
- 品质
- 主动性
- Just Talk To It - the no-bs Way of Agentic Engineering
- Quantifying infrastructure noise in agentic coding evals
- Infrastructure noise in evals
- Introducing Contextual Retrieval
- Contextual Retrieval
- Effective context engineering for AI agents
- Context engineering
- The “think” tool- Enabling Claude to stop and think in complex tool use situations
- Reasoning tools for agents
- Effective harnesses for long-running agents
- Multi-context window workflows
- Best Practices for Claude Code
- Context hygiene for agents
- Scaling Managed Agents-Decoupling the brain from the hands
- Meta-harness
- Beyond permission prompts- making Claude Code more secure and autonomous
- Sandboxing for agents
- How we built our multi-agent research system
- Multi-agent research systems
- Building effective agents
- Workflows vs agents
- Demystifying evals for AI agents
- Evaluation harness
- Claude Code auto mode- a safer way to skip permissions
- Permission delegation for agents
- A postmortem of three recent issues
- Inference infrastructure regressions
- Introducing advanced tool use on the Claude Developer Platform
- Programmatic tool calling
- Writing effective tools for agents — with agents
- Tool ergonomics for agents
- Harness design for long-running application development
- Generator-evaluator loop
- Building a C compiler with a team of parallel Claudes
- Agent teams
- Designing AI-resistant technical evaluations
- AI-resistant technical evaluations
- Code execution with MCP - Building more efficient agents
- Model Context Protocol (MCP)
- Eval awareness in Claude Opus 4.6’s BrowseComp performance
- Shipping at Inference-Speed
- Claude Opus 4.6
- BrowseComp
- Eval awareness
- Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
- Anthropic
- Claude 3.5 Sonnet
- SWE-bench Verified
- Agent scaffold
- Agentic coding evals
- Agentic engineering
- 导入来源:llm-wiki/raw/03_成长/品质
- 导入来源:llm-wiki/raw/03_成长/《不要害怕任何人和任何事》
- 导入来源:llm-wiki/raw/03_成长/张一鸣2016年演讲
- 导入来源:llm-wiki/raw/03_成长/孙宇晨为什么能这么成功?
- 导入来源:llm-wiki/raw/01_AI/国务院关于深入实施“人工智能+”行动的意见
- 导入来源:[llm-wiki/raw/peter blog/Just Talk To It - the no-bs Way of Agentic Engineering](/raw/peter blog/Just Talk To It - the no-bs Way of Agentic Engineering.md)
- 导入来源:[llm-wiki/raw/anthropic/Quantifying infrastructure noise in agentic coding evals](/raw/anthropic/Quantifying infrastructure noise in agentic coding evals.md)
- 导入来源:[llm-wiki/raw/anthropic/Introducing Contextual Retrieval](/raw/anthropic/Introducing Contextual Retrieval.md)
- 导入来源:[llm-wiki/raw/anthropic/Effective context engineering for AI agents](/raw/anthropic/Effective context engineering for AI agents.md)
- 导入来源:[llm-wiki/raw/anthropic/The “think” tool- Enabling Claude to stop and think in complex tool use situations](/raw/anthropic/The “think” tool- Enabling Claude to stop and think in complex tool use situations.md)
- 导入来源:[llm-wiki/raw/anthropic/Effective harnesses for long-running agents](/raw/anthropic/Effective harnesses for long-running agents.md)
- 导入来源:[llm-wiki/raw/anthropic/Best Practices for Claude Code](/raw/anthropic/Best Practices for Claude Code.md)
- 导入来源:[llm-wiki/raw/anthropic/Scaling Managed Agents-Decoupling the brain from the hands](/raw/anthropic/Scaling Managed Agents-Decoupling the brain from the hands.md)
- 导入来源:[llm-wiki/raw/anthropic/Beyond permission prompts- making Claude Code more secure and autonomous](/raw/anthropic/Beyond permission prompts- making Claude Code more secure and autonomous.md)
- 导入来源:[llm-wiki/raw/anthropic/How we built our multi-agent research system](/raw/anthropic/How we built our multi-agent research system.md)
- 导入来源:[llm-wiki/raw/anthropic/Building effective agents](/raw/anthropic/Building effective agents.md)
- 导入来源:[llm-wiki/raw/anthropic/Demystifying evals for AI agents](/raw/anthropic/Demystifying evals for AI agents.md)
- 导入来源:[llm-wiki/raw/anthropic/Claude Code auto mode- a safer way to skip permissions](/raw/anthropic/Claude Code auto mode- a safer way to skip permissions.md)
- 导入来源:[llm-wiki/raw/anthropic/A postmortem of three recent issues](/raw/anthropic/A postmortem of three recent issues.md)
- 导入来源:[llm-wiki/raw/anthropic/Introducing advanced tool use on the Claude Developer Platform](/raw/anthropic/Introducing advanced tool use on the Claude Developer Platform.md)
- 导入来源:[llm-wiki/raw/anthropic/Writing effective tools for agents — with agents](/raw/anthropic/Writing effective tools for agents — with agents.md)
- 导入来源:[llm-wiki/raw/anthropic/Harness design for long-running application development](/raw/anthropic/Harness design for long-running application development.md)
- 导入来源:[llm-wiki/raw/anthropic/Building a C compiler with a team of parallel Claudes](/raw/anthropic/Building a C compiler with a team of parallel Claudes.md)
- 导入来源:[llm-wiki/raw/anthropic/Designing AI-resistant technical evaluations](/raw/anthropic/Designing AI-resistant technical evaluations.md)
- 导入来源:[llm-wiki/raw/anthropic/Code execution with MCP - Building more efficient agents](/raw/anthropic/Code execution with MCP - Building more efficient agents.md)
- 参考翻译:通过 MCP 进行代码执行——构建更高效的智能体
- 导入来源:[llm-wiki/raw/peter blog/Shipping at Inference-Speed](/raw/peter blog/Shipping at Inference-Speed.md)
- 导入来源:[llm-wiki/raw/anthropic/Eval awareness in Claude Opus 4.6’s BrowseComp performance](/raw/anthropic/Eval awareness in Claude Opus 4.6’s BrowseComp performance.md)
- 参考翻译:Claude Opus 4 在 BrowseComp 上的评估意识
- 导入来源:[llm-wiki/raw/anthropic/Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet](/raw/anthropic/Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet.md)
- 参考翻译:用 Claude 3.5 Sonnet 提升 SWE-bench Verified 的标准