Anthropic
Anthropic
Section titled “Anthropic”- 12- — Constitutional AI: RL from AI Feedback 工程范式分析
- 07- — Claude 2 System Card — RLHF + Constitutional AI 的开端
- 01- — Sleeper Agents: 欺骗性对齐的工程范式分析
- 03- — Claude 3 Model Card: 首次多模型发布工程范式分析
- 05- — Scaling Monosemanticity: 可解释性里程碑的工程范式分析
- 06- — Claude 3.5 Sonnet/Opus System Card — 能力跃迁 + Agentic 转向
- 12- — Alignment Faking: 对齐伪装现象的工程范式分析
- 01- — Constitutional Classifiers: 安全分类器工程范式分析
- 03- — Circuit Tracing / On the Biology of a LLM — 可解释性新方法
- 05- — Claude 4 / Opus 4 / Sonnet 4 System Card — Hybrid Reasoning + ASL-3 首次部署
- 05- — Claude Opus 4.5/4.6/4.7/4.8 System Card 合集 — 快速迭代中的安全工程
- 06- — Claude Fable 5 & Mythos 5 System Card — 安全分层工程
- 06- — Claude Sonnet 5 — 最有 Agent 能力的 Sonnet