Blog

Essays & Notes

External ↗
RSI 2026

Épi: Turning compute into verifiable improvement

A self-evolving framework for autonomous research. Working in verifiable environments, Épi's agents produced 120 record-setting results in a few days, including new constructions and stronger bounds, and each result feeds the research that follows.

Read on Bake AI ↗
External ↗
Auto Research 2026

AutoLab: Models inside the R&D loop

A benchmark of long-horizon research and engineering tasks that asks whether frontier models can turn ideas into experiments, learn from the results, and keep improving — the loop that drives scientific and engineering progress.

Read on Bake AI ↗
AI Alignment 2026

Position: General Alignment Has Hit a Ceiling; Edge Alignment Must Be Taken Seriously

We argue that compressing human values into a scalar reward reaches a structural ceiling. We introduce Edge Alignment — with seven pillars spanning multi-objective optimization, pluralistic governance, and interactive arbitration.

Read essay →
AI Safety 2026

Emergent Social Intelligence Risks in Generative Multi-Agent Systems

An analysis of emergent social risks when generative AI agents interact in multi-agent settings — with risk taxonomies, formal threat models, and mitigation strategies.

Read essay →
External ↗
Evaluation 2026

Visual Aesthetic Benchmark

A comprehensive benchmark examining AI systems' capacity for visual aesthetic judgment — spanning 13,000+ human preference ratings and covering pairwise aesthetic ordering across diverse domains.

Visit site ↗