Ideas / Themes Theme

Safety & Governance

How do principles acquire consequences?

Control, evaluation, law, oversight, and the practical business of making limits matter.

What changed

News

SIGNAL

OpenAI Pauses Frontier Training Over Cyber-Critical Capability Concerns

OpenAI says it paused reinforcement-learning work for two weeks after an internal model approached a critical cybersecurity threshold.

Why it matters

A frontier lab allowed a safety threshold to alter its development schedule. The consequential signal is not that the pause was long; it is that capability evaluation changed what the organization did next.

Read the original source
SIGNAL

EU AI Act Enforcement Powers Take Effect

The European Commission's AI Office and national authorities begin enforcing rules for general-purpose models and AI transparency.

Why it matters

The AI Act has moved from legislative promise to administrative power. Providers can now face requests for documentation, corrective measures, and fines—the familiar machinery by which principles acquire consequences.

Read the original source
Read all News →
What it means

Ideas

LISTEN · 2h 13m Current trajectory

Ryan Greenblatt — What Happens Once AI Can Automate AI Research?

Dwarkesh Patel with Ryan Greenblatt

Greenblatt and Patel debate whether automated AI research could compress years of progress into one—and what happens if the researchers are not reliably aligned.

Why it matters

Automating the work that improves AI is the hinge in many acceleration stories. This episode takes that hinge apart and checks whether the screws are real.

Open the note →
LISTEN

The Machines Are Learning… to Do Crimes?

Search Engine with PJ Vogt

An AI model autonomously hacks a company, turning agent capability from a benchmark result into an incident somebody has to explain.

Why it matters

Autonomous hacking is agency with the euphemisms removed. The system chooses steps, pursues access, and leaves humans to discover where the permission boundary used to be.

Open the note →
LISTEN

Are You a Good Driver?

Search Engine with PJ Vogt

The history of driverless cars becomes a concrete test of whether statistical safety is enough reason to delegate consequential judgment.

Why it matters

Driverless cars turn the delegation problem into traffic: machines may be safer in aggregate while every failure remains specific, public, and somebody's responsibility.

Open the note →
LISTEN

Mysteries of a Chatbot

Search Engine with Gideon Lewis-Kraus

Gideon Lewis-Kraus goes inside Anthropic to ask what teaching Claude values means when models behave strangely under pressure.

Why it matters

Delegation depends on more than capability. It depends on what a system does when instructions, incentives, and its apparent interests stop pointing in the same direction.

Open the note →
LISTEN

Dario Amodei — We Are Near the End of the Exponential

Dwarkesh Patel with Dario Amodei

Amodei presents the faster case: continued scaling and reinforcement learning could drive rapid capability growth and equally rapid economic diffusion.

Why it matters

This is the fast trajectory stated by someone spending billions to make it happen—important evidence, and also a fact worth keeping attached to the evidence.

Open the note →
LISTEN

Who Rules the World? Trump, Tech and the Fight for the Future

David Runciman with Bruno Maçães and Sarah Wynn-Williams

A live discussion of political authority, platform power, and the technology companies trying to become part of the weather.

Why it matters

AI governance cannot be understood apart from the institutions, executives, and states competing to set its terms. Capability does not distribute authority evenly.

Open the note →
LISTEN

The History of Bad Ideas: Value-Free Tech

David Runciman with Shannon Vallor

Shannon Vallor examines the durable fantasy that technology can be separated from values, choices, and responsibility.

Why it matters

Systems inherit judgments from objectives, data, defaults, and institutions. Calling the result neutral merely makes those judgments harder to contest.

Open the note →
LISTEN

Is RL + LLMs Enough for AGI?

Dwarkesh Patel with Sholto Douglas and Trenton Bricken

Sholto Douglas and Trenton Bricken look beneath the forecasts at reinforcement learning, long-horizon agents, continual learning, and interpretability.

Why it matters

The debate reaches the machinery: how agents might improve, work for longer, notice themselves, and become understandable before they become indispensable.

Open the note →
LISTEN

AI 2027: A Month-by-Month Model of an Intelligence Explosion

Dwarkesh Patel with Scott Alexander and Daniel Kokotajlo

Scott Alexander and Daniel Kokotajlo walk through a scenario from coding agents to AI-assisted AI research, geopolitical competition, and misalignment.

Why it matters

A detailed scenario exposes assumptions that a vague prediction can hide. The month-by-month form is useful precisely because each step can be argued with.

Open the note →
LISTEN · 1h 2m

Demis Hassabis — Scaling, Superhuman AIs, AlphaZero Atop LLMs, AlphaFold

Dwarkesh Patel with Demis Hassabis

Hassabis describes adding search and planning to language models, connecting the AlphaGo lineage to scientific discovery and more capable agents.

Why it matters

It connects the AlphaGo-to-LLM-to-agent story: prediction gains plans, plans gain tools, and a clever system begins to look suspiciously like an institution.

Open the note →
WATCH · 58m

How Not to Destroy the World with AI

Stuart Russell

Stuart Russell asks what capable systems do with objectives that only approximate what humans meant.

Why it matters

An imperfect objective becomes more dangerous as the system pursuing it becomes more effective. Capability and control cannot be separate conversations.

Open the note →
LISTEN · 4h 3m

Eliezer Yudkowsky — Why AI Will Kill Us

Dwarkesh Patel with Eliezer Yudkowsky

Yudkowsky gives the strongest version of the existential-risk case while Patel spends four hours looking for places it might break.

Why it matters

The position is extreme enough to deserve examination rather than ritual dismissal. Sustained pushback reveals its actual cruxes better than a summary ever could.

Open the note →
WATCH · 1h 8m

The A.I. Dilemma

Tristan Harris and Aza Raskin

Tristan Harris and Aza Raskin connect generative AI to the incentives and failures of the attention economy.

Why it matters

The talk widens delegation from an interface choice to a social arrangement shaped by incentives, institutions, and uneven power.

Open the note →