
Regulatory sandboxes arrived with a seductive pitch. Give innovators a safe, bounded space to test new products, suspend a few rules, watch what happens, and then decide if the rulebook needs a rewrite. It’s a tidy image—a kind of legal laboratory where the state steps back just enough to let the future audition. But the image hides a knot of structural tensions that only become visible when you try to move from a handful of carefully chosen experiments to something that looks like a general-purpose governance tool. To see those tensions clearly, you have to look at what a sandbox actually does, and what it demands of the institutions that run one.
The Anatomy of a Sandbox
Strip it down, and a regulatory sandbox is a structured exemption. A firm applies to test a product, service, or business model that would otherwise break existing rules. If the regulator says yes, the firm gets a time-limited waiver, usually with strings attached: caps on customer numbers, mandatory disclosures, extra reporting, or specific consumer-protection backstops. In return, the regulator gets something it normally lacks—a real-time view of how a novel offering behaves in a live but carefully fenced-in market.
This isn’t deregulation. It’s a supervised, deliberate departure from the standard playbook. The sandbox doesn’t suspend liability; it rearranges it. The firm still answers for harms, and the regulator still answers for the integrity of the test. What shifts is the ex ante compliance burden. Instead of proving full conformity before launch, the firm proves it can manage risk inside the sandbox’s guardrails. The regulator, meanwhile, moves from gatekeeper to observer—though never completely, because deciding who gets in is itself a gatekeeping act.
The model was born in financial services. The U.K.’s Financial Conduct Authority launched the first formal sandbox in 2016. Since then, the idea has jumped sectors—energy, health, transport, data protection—and crossed borders, from Singapore to Arizona. Each adopter bends the template to fit its own legal architecture, but the basic choreography stays the same: apply, assess, test, exit, evaluate.
What Sandboxes Actually Produce
Advocates like to call sandboxes engines of innovation. The evidence points to something narrower. A 2019 review of the FCA’s early cohorts found that sandbox firms cut their time-to-market and had an easier time raising money. But the review also noted that the sandbox’s main value wasn’t regulatory relief as such. It was the signal. Getting into the sandbox worked as a credential, a stamp of preliminary regulatory approval that calmed investors and partners. The sandbox, in other words, functioned partly as a reputational device.
That’s not a trivial outcome. In sectors where regulatory uncertainty freezes investment, a credible signal can unlock capital. But it also means the sandbox’s benefits are tied to its exclusivity. If everyone gets a badge, the badge stops meaning anything. The signaling power depends on scarcity—on the regulator’s willingness to say no. And that sets up a tension between the sandbox as a learning tool and the sandbox as an endorsement machine.

The Scaling Problem Begins with Admission
If you want to understand why sandboxes resist scaling, start with the application process. A well-run sandbox demands case-by-case assessment. Each applicant proposes a different departure from the rules, a different risk profile, a different set of mitigants. The regulator has to evaluate not just the firm’s competence but the proportionality of the safeguards it’s offering. This is labor-intensive, expertise-intensive work. It doesn’t lend itself to automation or standardization, because the whole point is to handle what the standard rulebook can’t.
When a sandbox stays small—say, ten or twenty firms per cohort—the resource demands are manageable. A dedicated team can do deep due diligence, negotiate bespoke conditions, and keep close supervisory contact throughout the test. But try to scale that to hundreds of firms, and two things break. First, the quality of assessment degrades; the regulator starts making coarse judgments, which eats away at the sandbox’s legitimacy. Second, the supervisory relationship thins. The regulator can no longer watch each firm closely, so the sandbox becomes a lighter-touch regime by default, not by design.
This isn’t a failure of imagination. It’s a consequence of the sandbox’s defining feature: tailored oversight. Scale demands standardization, but standardization erases the very flexibility that makes a sandbox useful. You end up with something closer to a class waiver or a general permit—instruments that have their own logic and their own limits, but that aren’t sandboxes.
Legal Architecture and the Problem of Authority
Sandboxes also run into the hard edges of administrative law. In plenty of jurisdictions, regulators don’t have unlimited discretion to waive rules. Their powers are bounded by statute, and those statutes often require uniform application of regulations. A sandbox, by design, creates unequal treatment: one firm gets a waiver, another doesn’t. That inequality has to be legally justified, usually by pointing to the regulator’s mandate to promote innovation or competition. But the broader the sandbox becomes, the harder it is to defend those justifications against challenges from firms left outside.
Take a hypothetical. A fintech sandbox admits fifty firms in a year, all of which get relief from certain capital requirements. A traditional bank, stuck with the full weight of those requirements, argues the sandbox creates an unlevel playing field. The regulator has to show the sandbox serves a legitimate purpose and that the unequal treatment is proportionate. If the sandbox is small and experimental, that argument is easier to make. If the sandbox is large and semi-permanent, the argument weakens. The sandbox starts to look less like an experiment and more like a parallel regulatory track—one the legislature never authorized.
This isn’t a hypothetical everywhere. In some countries, sandboxes have been challenged on exactly these grounds, forcing regulators to anchor them more firmly in primary legislation. But legislative authorization brings its own constraints. Parliaments tend to set limits on scope, duration, and the types of rules that can be waived. Those limits are sensible from a rule-of-law perspective, but they also cap the sandbox’s reach. The sandbox becomes a niche instrument, not a general-purpose one.
Consumer Protection in a Bounded Space
Consumer protection is the sandbox’s most sensitive spot. The standard justification is that sandbox participants have to provide adequate safeguards—disclosure, redress mechanisms, compensation arrangements—so that consumers are no worse off than under the normal rules. In practice, “no worse off” is a hard standard to verify. Sandbox firms often serve early adopters who are more tolerant of risk, which can mask harms that would show up at scale. And the sandbox’s time-limited nature means long-term effects—data misuse, erosion of trust, subtle forms of lock-in—may not surface before the test ends.
Regulators deal with this by imposing exit plans: what happens to consumers when the sandbox period expires? If the firm can’t transition to full authorization, it has to wind down the service without stranding customers. But exit planning is only as strong as the firm’s solvency and the regulator’s enforcement capacity. A firm that fails during the sandbox may not have the resources to manage an orderly exit. The regulator then faces a choice: absorb the cost, leave consumers exposed, or quietly extend the sandbox to avoid a messy collapse. None of these options fits the sandbox’s original logic.

Learning That Doesn’t Travel
Maybe the deepest scaling limit is epistemic. A sandbox is supposed to generate knowledge that feeds back into rulemaking. The regulator learns what works, what fails, and what the existing rules missed. But the knowledge a sandbox produces is highly contextual. It depends on the specific firms admitted, the specific waivers granted, the specific market conditions during the test period. Generalizing from a handful of bespoke experiments to a sector-wide rule is a leap most regulatory lawyers and economists would hesitate to make.
This isn’t a problem if the sandbox stays small. A few dozen tests can yield useful insights about regulatory friction points, even if those insights aren’t statistically representative. But if the goal is to use sandboxes as a routine policy-development tool—to run hundreds of tests and then rewrite the rulebook—the epistemic gap widens. The regulator is no longer learning from exceptions; it’s trying to derive general rules from a collection of exceptions, each of which was designed to be exceptional. The logic inverts.
Some jurisdictions have tried to fix this by building formal feedback loops: sandbox findings feed into a dedicated innovation unit, which then proposes rule changes. But the unit still faces the same generalization problem. It can spot patterns, but it can’t easily tell the difference between patterns that reflect genuine market evolution and patterns that are artifacts of the sandbox’s own selection criteria. The sandbox teaches you about the firms that entered the sandbox, not necessarily about the market as a whole.
Institutional Capacity and the Quiet Drift
There’s also a subtler institutional dynamic at work. When a regulator launches a sandbox, it often creates a dedicated team with a distinct culture: more open to experimentation, more comfortable with uncertainty, more willing to engage with firms as partners rather than as regulatees. That culture is hard to maintain as the sandbox grows. The team has to expand, drawing in staff from other parts of the organization who may not share the same ethos. The sandbox’s internal processes become more bureaucratic, more risk-averse, more like the standard supervisory approach it was meant to complement.
This drift isn’t inevitable, but it’s common. It reflects a basic tension between the sandbox’s role as an exception to the rule and the organization’s default instinct to regularize exceptions. Over time, the sandbox accumulates its own rulebook: eligibility criteria, standard conditions, precedent-based decision-making. It becomes a mini-regime, complete with its own orthodoxies. The space for genuine experimentation narrows, not because anyone decided to narrow it, but because institutions absorb novelty and make it routine.
When Sandboxes Make Sense
None of this means sandboxes are useless. They fit a specific set of circumstances: when the regulatory barrier is clear but the risk is poorly understood; when the number of potential applicants is small; when the regulator has the legal authority to grant tailored waivers without creating systemic inequities; and when there’s a credible pathway from sandbox test to permanent rule change. In those conditions, a sandbox can be a precise, low-cost way to gather evidence and build regulatory competence in an emerging area.
But those conditions are rare. More often, the sandbox is deployed as a general-purpose response to “innovation,” without a clear theory of what regulatory problem it solves. It becomes a signal of the regulator’s modernity rather than a carefully bounded instrument. And when that happens, the scaling limits—admission bottlenecks, legal fragility, consumer-protection gaps, epistemic thinness, institutional drift—begin to compound. The sandbox doesn’t fail; it just becomes something else, something less coherent.
Frequently Asked Questions
What is the difference between a regulatory sandbox and a pilot program?
A pilot program typically tests a predefined policy intervention under controlled conditions, often with a randomized or quasi-experimental design. A regulatory sandbox, by contrast, tests a firm’s product or service under a temporary waiver of rules. The sandbox focuses on firm-level experimentation; the pilot focuses on policy-level evaluation. Both can generate evidence, but their legal foundations and evidentiary standards differ.
Why can’t sandboxes be automated to handle more applicants?
The core work of a sandbox—assessing novel risks, negotiating bespoke conditions, and maintaining close supervisory contact—requires discretionary judgment that resists codification. While parts of the application process can be streamlined, the substantive evaluation depends on context-specific expertise. Attempts to automate that evaluation risk reducing the sandbox to a checklist, which undermines its purpose and may expose the regulator to legal challenge.
Do sandboxes weaken consumer protection?
Not necessarily, but they change its form. In a sandbox, consumer protection shifts from ex ante compliance with prescriptive rules to ex post reliance on tailored safeguards and supervisory oversight. Whether this shift weakens or strengthens protection depends on the quality of the safeguards, the regulator’s monitoring capacity, and the firm’s incentives. The risk is that the safeguards look sound on paper but prove fragile under stress, especially if the sandbox scales beyond the regulator’s ability to monitor closely.
Can sandboxes work outside financial services?
They can, but the challenges multiply. Financial regulators often have broad statutory mandates that allow for waivers and experimentation. In sectors like health or energy, the legal framework may be more rigid, and the risks of failure—physical harm, environmental damage—are harder to contain within a sandbox’s boundaries. The sandbox model also assumes that the regulator has sufficient technical expertise to evaluate the innovation, which may not hold in highly specialized domains.
The Unanswered Question of Legitimacy
Underneath the practical difficulties lies a deeper question: who authorized the sandbox? In a democratic regulatory system, rules aren’t just technical instruments; they’re expressions of public values, negotiated through legislative and administrative processes. When a regulator creates a sandbox, it’s effectively creating a parallel track that suspends some of those values for a select group of participants. That may be defensible as a limited experiment, but it becomes harder to defend as the sandbox grows and the suspensions become routine.
This isn’t an argument against sandboxes. It’s an argument for treating them as what they are: narrow, temporary, and exceptional instruments that require clear legal authorization, rigorous evaluation, and a predefined exit—either toward permanent rule changes or toward closure. The temptation to scale them is understandable, because they appear to solve a genuine problem: the mismatch between the pace of innovation and the pace of rulemaking. But scaling a sandbox doesn’t solve that mismatch; it just creates a parallel system with weaker accountability. The harder, more necessary work is to build regulatory frameworks that are adaptive by design, not by exception.
Recent Comments