“Frontier AI” is the term AI labs, governments, and safety researchers use for the most capable general-purpose AI systems that exist at any given moment — not one company’s product, but whichever models currently sit at the leading edge of what AI can do. Because that edge keeps moving, “frontier” is a relative label: a system counts as frontier today and becomes ordinary within a year or two, once newer models pass it.

What Counts as “Frontier”

There is no single legal definition of frontier AI. The most widely cited one comes from the Frontier Model Forum (FMF), an industry body founded in 2023 by Anthropic, Google, Microsoft, and OpenAI. It defines a frontier model as a general-purpose system that outperforms — on standard benchmarks or on assessments of high-risk capabilities — every model that has already been in wide use for at least 12 months. In practice, that means a model only earns the label by beating what came before it; the bar resets as each new generation ships.

The term entered wider policy use at the 2023 AI Safety Summit at Bletchley Park, where the UK government described frontier AI as “highly capable general-purpose AI models that can perform a wide variety of tasks and match or exceed the capabilities present in today’s most advanced models.” That framing is deliberately broad: it covers not just chatbots but any general-purpose system — including ones later wrapped into autonomous AI agents with tool use — capable of unexpected, emergent behavior as it scales.

How Governments Draw the Line

Regulators have tried to turn that fuzzy idea into something enforceable, usually by measuring the computing power used to train a model. The EU AI Act treats a general-purpose model as posing “systemic risk” — triggering extra transparency, testing, and incident-reporting duties — once its training run exceeds roughly 10^25 floating-point operations (FLOPs), with regulators able to add specific models below that line if evidence warrants it. The UK’s AI Safety Institute instead evaluates individual frontier models directly, through voluntary pre- and post-deployment testing agreements with the labs that build them. The United States has experimented with its own version too, screening which frontier systems get early government access before public release.

None of these thresholds are universal, and a model can be “frontier” under one definition while falling short of another. What they share is the underlying goal: singling out the small set of systems capable enough that mistakes or misuse could cause outsized harm, so that they get scrutiny ordinary software doesn’t.

Why It Matters

Frontier models matter because they are where AI’s biggest capability jumps — and its biggest risks — tend to show up first. A model at the frontier can pick up abilities its developers never explicitly trained for, simply by being bigger and trained on more data and compute. That unpredictability is why several labs now publish Responsible Scaling Policies: internal commitments to pause or add safeguards once a model’s capabilities cross specific risk thresholds, rather than waiting for regulation to catch up.

It’s also why the pace of frontier development has become a policy fight in its own right, not just a technical one. The labs building these systems compete fiercely to ship the next frontier model first, and that competitive pressure makes it hard for any single company to slow down unilaterally — even when its own staff have doubts. Governing that race, rather than any single model, is now central to how AI safety policy is framed.

In the News

That tension is exactly what drove more than 1,200 employees at frontier labs to act: in a joint petition, workers from companies including OpenAI, Anthropic, and Google DeepMind asked the US government to help build international tools to deliberately pace frontier AI development — not to halt it, but to keep oversight from falling permanently behind the technology it’s meant to govern.