Rules for scale.
Standards for nuance.
Exceptions for failure.
Principles for the unknown.
You cannot remove indeterminacy.
You can only decide where it lives.

There is a deceptively simple problem hiding inside almost every system of governance.
You want to say:
Drive safely.
But that is not a very useful law.
So you make it more concrete:
The speed limit is 55 mph.
Now you have a rule that people can actually follow. Police can enforce it. Courts can adjudicate it. Drivers know what is expected.
But you have also created a problem.
A person can drive at 70 mph on an empty, dry highway in a modern car, with excellent visibility and no one else around, and be driving quite safely. Meanwhile, someone can drive at 55 mph on an icy road in dense fog and be driving dangerously.
The rule is therefore not the same thing as the reason for the rule.
This looks like a small problem about traffic tickets. It turns out to be a deep problem about how we construct rules for humans—and potentially one of the most useful ways to think about governing LLMs.
I call it the Montana Problem.
Montana provides an unusually good story because it actually tried the obvious alternative.
Instead of relying entirely on fixed numerical speed limits, Montana's law used a general rule requiring motorists to drive at a speed that was "reasonable and proper under the conditions existing" and to take account of factors such as traffic, the vehicle, the road, visibility, and conditions. In State v. Stanko in 1998, the Montana Supreme Court held that the speed component of that rule was unconstitutionally vague. The court's reasoning was strikingly practical: motorists needed to know what conduct was prohibited, and the relevant judgments about sight distance, stopping distance, road geometry and so forth were not something that should be left to ad hoc judgments by individual drivers or police officers.
Montana therefore illustrates both sides of the problem.
A principle such as:
Drive safely.
is too vague.
A standard such as:
Drive at a reasonable and prudent speed given the circumstances.
is more context-sensitive, but can produce inconsistent judgments.
A hard rule such as:
Maximum speed: 55 mph.
is much easier to apply, but inevitably misclassifies some cases.
That is the puzzle.
And it is not peculiar to Montana. It is a classic problem in legal theory.
Louis Kaplow's famous 1992 paper, Rules Versus Standards: An Economic Analysis, gives perhaps the cleanest way of thinking about this.
Rules and standards don't merely differ in how precise they are. They move costs around.
A precise rule is expensive to formulate but cheap to apply.
A standard is cheap to formulate but expensive to apply consistently.
Kaplow's formulation is remarkably relevant to software and AI:
Where is it cheaper and better to resolve uncertainty: when writing the law, or when applying it?
That is the real design question.
Suppose a government wants to regulate 100 million driving decisions.
It could make every driver and police officer perform a contextual safety analysis every time. That gives you richer decisions, but the reasoning cost is enormous, and different people will make different judgments.
Or it can resolve much of the uncertainty in advance:
55 mph.
The government has paid the cost of reasoning once, when setting the rule, and then gets extremely cheap enforcement thereafter.
That is the economic logic of the bright-line rule. Kaplow explicitly treats the choice between rules and standards in terms of creation costs, interpretation costs, adjudication costs, the level of detail in legal commands, and the over- and under-inclusiveness that results.
The interesting point is not simply:
Rules need exceptions.
That is true but shallow.
The deeper point is:
Indeterminacy cannot be eliminated. You choose where it lives.
You can put uncertainty in the legislation.
You can put it in the police officer.
You can put it in the judge.
You can put it in an exception process.
You can put it in an administrative agency.
But somebody, somewhere, eventually has to decide what the abstract objective means in a particular situation.
Montana put a lot of that uncertainty into the person applying the standard.
A fixed speed limit puts much more of it into the person who designed the rule.
That is the fundamental tradeoff.
This is where Frederick Schauer's work becomes particularly important.
Schauer's Playing by the Rules asks a question that sounds almost perverse:
Why deliberately use a rule when you know the rule will sometimes produce the substantively wrong result?
Because that loss of contextual accuracy can buy something valuable.
A rule can constrain the person applying it.
It can reduce discretion.
It can make decisions predictable.
It can reduce inconsistency.
It can make compliance cheap.
It can prevent the person with authority from simply re-deciding every case according to their own view of the underlying purpose.
Rules therefore sometimes deliberately sacrifice perfect substantive accuracy for predictability and control. Recent scholarship summarizing Schauer's position emphasizes precisely this role of rules in reducing information asymmetries between rule-makers and rule-appliers and accepting some over- and under-inclusiveness as the price.
The 55-mph rule consequently isn't necessarily stupid because it catches a safe driver.
It may be rational to say:
We know this rule is imperfect. We prefer its imperfections to having millions of individual decisions made by millions of different people.
That is an important distinction.
Consider two versions of the same law.
In one system, a government has to employ a million traffic officers to make individualized judgments.
In another, it has a cheap system that can reliably assess the relevant circumstances automatically.
The optimal legal architecture could be different.
This is one of the most interesting implications for LLM governance.
LLMs provide scaled-up reasoning capability.
That does not magically make contextual judgment correct, but it potentially changes its cost.
Historically, a government may have said:
"We cannot have a human expert reason through every one of these cases."
So it creates a crude rule.
An LLM potentially makes a different architecture feasible:
"We can have a machine perform a contextual assessment in a very large fraction of cases."
There is already research moving in this direction. A 2023 study by John Nay used thousands of labels derived from U.S. court opinions to evaluate whether LLMs could understand legal standards. In that study, the strongest model tested reached 78% accuracy on the fiduciary-duty dataset, compared with 73% for an earlier model and 27% for GPT-3-era performance.
That does not establish that LLMs are ready to replace judges or police officers. But it demonstrates the basic proposition that motivated the experiment: language models can perform meaningful standard-like legal reasoning, and their capability has been improving.
That changes the economics of the rules-versus-standards problem.
The old question was:
Can we afford contextual judgment at this scale?
The emerging question is:
Given cheap machine reasoning, where should we now put the boundary between rules and standards?
Real jurisdictions have not settled on one universal architecture.
They have placed the indeterminacy in different places.
Germany provides a particularly clean example.
Section 3 of the German Road Traffic Regulations says that drivers must control their vehicles and adapt their speed to road, traffic, visibility and weather conditions, as well as their own abilities and the characteristics of the vehicle and load. At the same time, the regulation provides concrete maximum speeds.
So Germany does not attempt to make the numerical rule synonymous with safety.
It effectively has:
Standard: adapt your speed to circumstances.
and
Rule: certain maximum speeds apply.
These answer different questions.
The UK's Highway Code makes the distinction almost embarrassingly clear:
The speed limit is the absolute maximum; it does not mean that it is safe to drive at that speed in all conditions.
It separately tells drivers to reduce speed for hazards, pedestrians, weather, darkness and road conditions. It also says that drivers must not drive dangerously or without due care.
The UK therefore essentially says:
The number is a ceiling. It is not a complete theory of safe driving.
That is an elegant solution to the conceptual confusion.
Singapore's Road Traffic Act is even more interesting.
Section 63 makes exceeding an applicable speed limit an offence. Separately, section 64 prohibits driving recklessly or at a speed or in a manner dangerous to the public, explicitly requiring consideration of the circumstances, including the nature and condition of the road and the volume of traffic.
This creates a useful separation:
Did you exceed the limit?
is one question.
Was your driving dangerous?
is another.
Singapore does not have to make the first question carry the entire burden of the second.
Its enforcement system also distinguishes degrees of speeding rather than treating every excess identically. Penalties were strengthened from January 2026, with increased demerit points and composition sums; Singapore's Ministry of Home Affairs reported nearly 120,000 speeding violations in the first half of 2025.
This is an important design pattern:
Keep the cheap rule for the common case. Use a richer legal standard for the qualitatively different case.
Texas offers another architecture.
In a 2020 Court of Criminal Appeals case, the court considered a prosecution under a standard saying that speeding is unlawful when the speed is greater than is reasonable and prudent under the circumstances. The posted speed limit could serve as prima facie evidence, but it was not an automatic determination that the speed was unreasonable.
That is a fascinating middle ground.
The rule becomes:
Start here.
rather than:
The inquiry ends here.
The numerical rule remains extremely useful, but the legal system retains a contextual route when the ordinary presumption is inappropriate.
Indian law provides a useful contrast.
Section 112 of the Motor Vehicles Act establishes speed limits and section 183 provides the offence for contravening them. Indian courts have generally treated exceeding the prescribed limit as a distinct offence rather than inviting a driver to re-litigate whether the particular situation was sufficiently safe.
A 2025 decision in State v. Rajiv Kumar Singh is unusually explicit. The defendant argued that the road was empty, that there was no immediate danger, and that speed limits should depend on the circumstances of the particular day and time. The court rejected that argument, treating the speeding violation as one in which mens rea was irrelevant and saying that setting the limits based on circumstances was a policy matter for the responsible governmental authorities. It also rejected an argument that safer modern vehicles made the fixed limits unnecessary.
This is the bright-line philosophy in its clearest form:
Don't make the driver and police officer re-decide the policy every time.
If the limit is wrong, change the limit.
That is essentially put the indeterminacy upstream.
These jurisdictions give us a useful spectrum:
LOTS OF EX POST JUDGMENT
│
│ Montana
│ "reasonable and prudent"
│
│ Texas
│ rule as presumption
│
│ Germany / UK / Singapore
│ rule + independent standard
│
│ India / strict bright-line regimes
│ numerical rule
│
▼
LOTS OF EX ANTE SPECIFICATION
The point is not that Germany is "better than India" or that Texas has discovered the perfect system.
The point is that these are different engineering choices about where to resolve uncertainty.
And scale matters.
When a rule must govern millions of decisions every day, the value of a cheap, predictable rule rises dramatically.
Lon Fuller adds another dimension.
A governing norm has to be capable of actually guiding conduct. If it is impossibly vague, contradictory, retroactive, constantly changing, or applied inconsistently with what it says, it ceases to function properly as law.
This is exactly what makes Montana so memorable.
"Reasonable and prudent" sounds intelligent.
But if a driver cannot predict what conduct will lead to a penalty, the rule is not doing the basic job of a behavioral rule.
This matters enormously for AI.
An LLM policy can be philosophically beautiful and still be useless if the agent cannot reliably determine what it requires.
A governing norm has to be capable of guiding behavior.
That sounds obvious, but it places an important constraint on "just give the model principles."
There is one more move that becomes important when rules genuinely run out.
Ronald Dworkin's famous critique of simple rule-based pictures of law emphasizes the role of principles. Legal reasoning does not consist solely of mechanically applying rules; principles can have a role in difficult cases.
For LLM governance, the practical lesson is straightforward:
When a lower-level rule genuinely fails to resolve a novel or exceptional case, the system needs a way to reason upward rather than simply guessing another rule.
This is different from saying:
"Ignore the rules whenever you have a good reason."
That destroys the value of rules.
The architecture should instead be:
Apply the rule. If an explicitly recognized exception or standard is implicated, reassess. If there is a genuine conflict or novel case, reason from the governing principle.
Most cases should stop early.
Only unusual cases should become expensive.
This is where the legal theory becomes directly useful for AI.
Suppose we are governing an agent.
We might write:
HIGH-ORDER PRINCIPLES
What are we ultimately protecting?
│
▼
STANDARDS
How should this class of situations
ordinarily be evaluated?
│
▼
HARD RULES
What should happen in normal cases?
│
▼
EXCEPTIONS / REBUTTALS
What facts justify reopening the case?
│
▼
PARTICULAR DECISION
This gives the model an escalation path.
At level one:
Rule: Do X.
Cheap. Fast. Consistent.
If challenged:
Standard: Does the context satisfy the conditions under which the rule should be relaxed?
More expensive. More nuanced.
If that still produces an unacceptable or genuinely novel result:
Principle: What higher-order objective was this entire structure intended to serve?
Now the model can reason from first principles.
This is a much better architecture than either extreme:
"Just give the agent rules."
or:
"Just give the agent principles and let it figure everything out."
The first creates brittleness.
The second creates uncontrolled discretion.
There is a subtle but important point here.
If the LLM is allowed to invoke the higher-order principle whenever it finds the rule inconvenient, you have recreated Montana.
Everything becomes:
"Well, in context..."
Now the rule no longer constrains behavior.
The value of layered governance comes from making escalation progressively rarer and more expensive.
Think of it as:
RULE
│
│ ordinary case
▼
DECISION
RULE
│
│ challenged
▼
EXCEPTION / STANDARD
│
│ justified?
├──── no ────► DECISION
│
▼
REASSESS
RULE + STANDARD
│
│ genuinely novel conflict
▼
HIGHER-ORDER PRINCIPLE
│
▼
REASONED DECISION
The rule is therefore not an obsolete layer that the intelligent model should bypass.
The rule is the optimization.
It handles the enormous number of ordinary cases cheaply.
Standards handle the smaller number of cases where context matters.
Principles handle the tiny number where the existing machinery itself breaks down.
There is an important qualification.
It would be a mistake to conclude:
"LLMs are intelligent, therefore we should replace rules with standards."
The legal lesson is precisely the opposite.
The rules-versus-standards problem remains. What changes is the cost structure.
If contextual reasoning becomes dramatically cheaper, standards become affordable in more situations.
But rules still buy:
And LLMs bring their own problems.
A contextual judgment made cheaply can still be inconsistent.
A model can rationalize an answer after the fact.
A standard can be interpreted differently across models or model versions.
A model can effectively invent millions of tiny distinctions that were never intended by the policy.
And this last point is now appearing explicitly in legal scholarship.
A 2026 paper in the Israel Law Review, "Bending the Rules: On Large Language Models and Content Moderation," applies rules-versus-standards theory directly to LLM-based moderation.
Its authors identify a paradox: platforms may formulate increasingly detailed, rule-oriented policies while LLM enforcement behaves more like a contextual standard. They introduce the phrase "rules by the millions" to describe how LLMs can effectively operate through vast networks of context-sensitive micro-rules.
This is remarkably close to the problem we have been describing.
Traditional law has:
Rule → application → exception.
An LLM may effectively produce:
Policy → enormous number of learned contextual distinctions → decision.
That can be extremely powerful.
It can also make the actual governing logic opaque.
The paper's central concern is therefore not simply whether the LLM is "following the rules." It is whether humans can understand, audit and govern the enormous amount of implicit decision logic produced by the model.
That is exactly why the legal architecture matters.
Instead of asking:
"What rules should I give the model?"
ask:
"At what level should each uncertainty be resolved?"
For every policy decision, ask:
Anything that is:
Anything where:
Things that occur rarely enough that you don't want to complicate the normal path, but are important enough that mechanical application would be unacceptable.
Cases that the policy writer genuinely cannot enumerate in advance.
This is not merely a hierarchy of prose.
It is a hierarchy of computational cost and discretion.
Return to the driver.
A sensible legal system might say:
Rule: 55 mph.
Then:
Standard: speed must also be appropriate to road, traffic and weather conditions.
Then:
Exceptions: emergency circumstances and specifically defined situations.
And finally:
Principle: the entire regime exists to promote road safety and orderly road use.
Notice what happened.
The system did not try to encode every possible driving situation into the rule.
It also did not abandon the rule and make every traffic stop a philosophical inquiry.
It created layers with different jobs.
That is exactly what an LLM policy can do.
This is where I think the most interesting idea emerges.
Historically, human institutions often had to simplify rules because human judgment does not scale cheaply.
A government with ten million drivers cannot afford ten million expert legal opinions every morning.
An organization with millions of users cannot afford a lawyer to examine every content-moderation decision.
An enterprise cannot have a senior engineer personally adjudicate every agent decision.
LLMs potentially change that constraint.
They can perform contextual interpretation at a much larger scale than humans can.
That does not make their judgments infallible. It means that the economic argument for pushing everything downward into crude rules becomes weaker.
Perhaps we can afford:
India-style cheap rules for ordinary cases
plus
UK/Germany/Singapore-style contextual standards
plus
human-like escalation to higher-order principles for genuinely novel cases
at a scale that previously required enormous administrative machinery.
The important phrase is "afford."
AI does not tell us that contextual judgment is always better.
It changes the answer to:
How much contextual judgment can we economically afford?
Montana initially looks like a story about a foolish driver being unable to understand "reasonable and prudent."
It is actually a story about where a society chooses to locate uncertainty.
A pure principle:
Drive safely.
is too vague.
A pure standard:
Drive reasonably and prudently.
moves too much discretion into individual application.
A pure rule:
55 mph.
is wonderfully cheap and predictable, but inevitably crude.
The mature solution is usually somewhere in between.
And the answer is not necessarily one particular point on the continuum.
It depends on:
frequency × scale × cost of judgment × cost of error × need for consistency × value of contextual accuracy.
That is the real lesson from Kaplow.
Schauer adds:
Don't underestimate what rules buy you.
Hart adds:
Some indeterminacy is unavoidable.
Fuller adds:
A governing norm has to be capable of guiding behavior.
Dworkin adds:
When rules genuinely fail, higher-order principles matter.
And actual jurisdictions demonstrate that these ideas can be implemented in very different ways.
Putting all of this together, I would formulate the lesson for people writing agent policies like this:
Indeterminacy cannot be eliminated; decide where to put it.
Put high-frequency, low-context decisions into hard rules.
Use standards where contextual accuracy is worth the additional reasoning cost.
Use narrow exceptions to prevent predictable pathological applications.
Escalate genuinely novel or irreconcilable cases to higher-order principles.
Do not allow higher-order reasoning to casually override ordinary rules, or you have recreated the Montana problem.
The resulting system optimizes different things at different levels:
| Layer | What it optimizes |
|---|---|
| Hard rules | Speed, consistency, cost, predictability |
| Exceptions | Handling known failure modes |
| Standards | Contextual accuracy and fairness |
| Higher-order principles | Novelty, coherence, resolution of genuine conflicts |
Most decisions should be cheap.
Some should be thoughtful.
Very few should require philosophical reasoning.
That is the point.
The traditional question was:
Rules or standards?
The more useful question for AI is:
What should be resolved before the agent acts, what should be resolved while the agent acts, and what should be escalated when the existing specification breaks down?
That is a much richer design space.
LLMs do not eliminate law's old problem.
They give us a machine capable of participating in all the layers of the legal reasoning process—and potentially make some of the expensive layers cheap enough to use at scale.
The challenge for AI governance is therefore not to choose between rules and intelligence.
It is to architect the relationship between them.
And when you're unsure how to do that, remember Montana.
A rule that tries to decide everything becomes stupid.
A standard that asks someone to decide everything becomes unpredictable.
Good governance knows which decisions should be made once, which should be made repeatedly, and which should be reopened only when something genuinely unusual happens.