In the world of financial and technology policy, the regulatory sandbox has been embraced with a fervor that borders on the evangelical. Since the United Kingdom launched the first formal version in 2016, the model has spread to dozens of jurisdictions, from Singapore to Arizona. The pitch is elegant: give innovative firms a temporary, supervised space where they can test new products without immediately shouldering the entire weight of the regulatory apparatus. For a sector where the cost of a license can dwarf the cost of building the software, this is a genuinely meaningful offer. But after nearly a decade of global experimentation, a more sober question has surfaced. Sandboxes are very good at producing pilots. They are far less reliable at producing markets.
I have spent the last several years studying how these frameworks interact with the actual mechanics of scaling a regulated service. The pattern is remarkably consistent across jurisdictions. A sandbox cohort will yield a handful of promising prototypes, a few modest regulatory tweaks, and a great deal of positive press. What it rarely yields is a clear, well-lit path from a supervised test with 500 users to a fully authorized deployment serving 500,000. The reasons for this are not incidental. They are structural, and they deserve a closer look than the typical policy brief provides.
The Architecture of a Sandbox
To understand the scaling problem, one must first understand what a sandbox actually does. At its core, it is a legal instrument that allows a firm to operate under a temporary waiver or modification of specific rules. The firm must apply, making the case that its product is genuinely novel, that it offers a consumer benefit, and that it cannot easily be tested within the existing rulebook. If accepted, the firm receives a limited authorization—often for six to twelve months—with strict guardrails: a cap on the number of clients, a ceiling on transaction volumes, and mandatory disclosures informing consumers they are part of an experiment.
This design is deliberate. It reflects a careful balance between encouraging novelty and protecting the public. The sandbox is, in essence, a controlled experiment. The regulator learns how a new business model interacts with the existing legal framework, and the firm learns whether its product can function under something resembling real-world conditions. The trouble is that the conditions inside the sandbox are not, in fact, very much like the real world at all.

The Scaling Gap
The central tension is straightforward: a sandbox is built to reduce risk, but scaling a business demands taking risks. When a firm exits a sandbox, it must move from a bespoke, closely monitored arrangement to the full regulatory regime. This transition is often called a “graduation,” but the metaphor is misleading. In practice, it feels more like leaving a sheltered workshop and being told to run a marathon the next morning.
Consider the case of a peer-to-peer lending platform that tested its model in a sandbox with a cap of 200 borrowers and a requirement that all lenders be accredited. The test went well: default rates were low, the technology held up, and consumers reported satisfaction. But when the firm sought full authorization, it collided with a completely different set of demands: capital adequacy ratios, anti-money laundering reporting systems, data protection audits, and a compliance department capable of handling thousands of transactions. The sandbox had answered the question “Does this product work?” but not the question “Can this business survive regulation?”
This is not a failure of the sandbox itself. It is a category error in how we measure success. A sandbox is a diagnostic tool, not a treatment plan. It reveals whether a regulatory barrier exists and whether a particular exemption can remove it. It does not, and cannot, reveal whether the firm can build the operational infrastructure to comply with the full rulebook once the exemption is lifted. That requires a different set of resources: capital, legal expertise, compliance staffing, and time. Most sandbox graduates lack at least two of these.
The Data Problem
Another structural limitation concerns data. Regulators often promote the sandbox as a learning mechanism—a way to gather evidence about new technologies before writing permanent rules. But the evidence generated by a sandbox test is, by design, thin. A sample of 200 users over six months tells you almost nothing about systemic risk, consumer harm at scale, or how the product behaves during a market downturn. It is a snapshot taken in a padded room. Extrapolating from that snapshot to a population of millions is not just ambitious; it is methodologically unsound.
This creates a paradox. Regulators want data before they amend rules, but the sandbox cannot produce data of sufficient quality to justify major rule changes. The result is often a policy limbo: the sandbox ends, the firm is told to apply for a full license, and the regulator promises to consider a rule change “in due course.” The firm, having burned through its seed funding during the sandbox period, cannot afford to wait. The innovation dies not because it failed, but because it succeeded in a system that had no next step.

Institutional Memory and Regulatory Capacity
Scaling sandbox innovations also depends on the regulator’s own capacity to learn and adapt. A sandbox is not just a test of the firm; it is a test of the regulator’s ability to process new information and translate it into rulemaking. This is where the model often breaks down. In many agencies, the sandbox team is a small, specialized unit that operates separately from the policy and supervision divisions. The insights generated during a test may never reach the people who write the rules or approve full licenses.
I have observed cases where a firm exits a sandbox with a positive evaluation, only to find that the licensing department has no record of the sandbox’s findings. The firm is treated as a new applicant, subject to the same requirements as any other. The sandbox experience becomes irrelevant. This is not malice; it is a failure of institutional knowledge transfer. The sandbox team and the licensing team inhabit different organizational silos, with different mandates and different timelines. Bridging that gap requires deliberate process design, which many regulators have not yet undertaken.
The Human Factor
There is also a subtler problem: the sandbox creates a relationship between the regulator and the firm that is unusually collaborative. Regulators assigned to sandbox cases often become advocates for the innovation they are supervising. They develop a deep understanding of the business model and a stake in its success. But when the sandbox ends, that relationship ends. The firm is handed off to a supervision team that has no prior connection to the case and may view the innovation with skepticism. The collaborative dynamic is replaced by a compliance dynamic, and the firm is often unprepared for the shift.
This is not an argument against sandboxes. It is an argument for designing them with the full lifecycle in mind. A sandbox that ends at graduation is incomplete. There must be a structured handover process, a clear pathway to full authorization, and a commitment from the regulator to use the sandbox evidence in rulemaking. Without these elements, the sandbox is a display case, not a pipeline.
Why Scaling Is Hard: A Structural View
To understand why sandbox innovations rarely scale, we need to look at the economics of regulation itself. Regulation is, at its core, a set of fixed costs imposed on firms. Compliance departments, reporting systems, legal reviews—these are all overhead that do not vary much with the number of customers. A firm with 500 users and a firm with 500,000 users may face similar compliance costs. This means that small firms, the kind that populate sandboxes, are disproportionately burdened by regulation. The sandbox temporarily relieves that burden, but it does not eliminate it. When the sandbox ends, the burden returns, and the firm’s unit economics often collapse.
This is not a problem that sandboxes can solve on their own. It requires a broader rethinking of proportional regulation—rules that scale with the size and risk profile of the firm. Some jurisdictions have begun to explore this, creating tiered licensing frameworks that impose lighter requirements on smaller players. But these efforts are still nascent, and they face resistance from incumbent firms and consumer advocates who worry about a race to the bottom. The sandbox, in this context, is a useful tool for testing proportionality, but it cannot substitute for the political work of building consensus around a new regulatory model.

What Works: Lessons from the Field
Despite these limitations, some sandboxes have produced lasting impact. The common thread, in almost every case, is that the sandbox was not treated as a standalone program but as part of a broader innovation strategy. The UK’s Financial Conduct Authority, for example, paired its sandbox with a “regulatory nursery” that provided extended support for graduating firms. Singapore’s Monetary Authority created a “Sandbox Express” for low-risk innovations, reducing the time and cost of entry. These additions acknowledge that the sandbox is a starting point, not a finish line.
Another factor is the type of innovation being tested. Sandboxes work best for products that are genuinely new and for which the regulatory barrier is a specific, identifiable rule. They work poorly for business models that require ongoing regulatory discretion, such as those involving fiduciary duties or complex suitability determinations. A robo-advisor, for instance, can test its algorithm in a sandbox, but the real regulatory challenge—ensuring that the advice is suitable for each individual client—cannot be simulated with a small, self-selected sample.
The Role of the Regulated Entity
It is also worth noting that the most successful sandbox graduates are often not startups at all, but subsidiaries of established financial institutions. These firms have the capital, compliance infrastructure, and regulatory relationships to survive the post-sandbox transition. The sandbox, for them, is a way to test a new product without risking their existing license. This is a perfectly legitimate use of the tool, but it is not the use case that most sandbox advocates emphasize. The narrative of the sandbox as a launchpad for disruptive startups is, in practice, more aspiration than reality.
Frequently Asked Questions
What is the primary purpose of a regulatory sandbox?
A regulatory sandbox is designed to allow firms to test innovative products, services, or business models in a controlled environment with temporary regulatory relief. The goal is to enable learning for both the firm and the regulator, identifying where existing rules may need adjustment without exposing consumers to undue risk.
Why do so few sandbox firms achieve full market authorization?
The transition from a sandbox to full authorization often reveals a mismatch between the firm’s operational capacity and the demands of the complete regulatory framework. Sandboxes typically waive only a subset of rules, and firms may lack the capital, compliance infrastructure, or institutional support to meet the remaining requirements at scale.
Can sandboxes be redesigned to improve scaling outcomes?
Yes, but it requires treating the sandbox as one phase of a longer innovation pathway. Effective reforms include structured handover processes between sandbox and licensing teams, tiered regulation that matches requirements to firm size and risk, and regulatory commitments to act on sandbox evidence within a defined timeframe.
Are sandboxes still worth pursuing despite these limitations?
Absolutely. Sandboxes provide valuable insights into emerging technologies and can help regulators identify outdated or disproportionate rules. The key is to set realistic expectations: a sandbox is a diagnostic tool, not a policy solution. Its value lies in the questions it raises, not just the firms it graduates.
The Path Forward
If sandboxes are to fulfill their promise, policymakers must resist the temptation to treat them as a cure-all. The hard work of regulatory reform happens outside the sandbox, in the tedious processes of rulemaking, legislative amendment, and international coordination. A sandbox can show that a rule is obsolete; it cannot rewrite the rule. That requires political will, technical expertise, and a willingness to confront the distributional consequences of change—who wins, who loses, and who pays.
There is also a need for more rigorous evaluation. Too many sandbox reports read like marketing brochures, highlighting success stories without quantifying failure rates or analyzing the reasons for attrition. A mature sandbox policy would include systematic follow-up with all participants, tracking their outcomes over three to five years and publishing the results. This would allow for evidence-based adjustments to the sandbox design and, more importantly, to the underlying regulatory framework.
Finally, we should be honest about the limits of the sandbox metaphor. A child’s sandbox is a place of unstructured play, where the stakes are low and the boundaries are clear. A regulatory sandbox is neither unstructured nor low-stakes. It is a carefully negotiated legal arrangement with real consequences for firms, consumers, and the integrity of the market. Treating it as a laboratory for policy learning, rather than a playground for innovation, would go a long way toward aligning expectations with outcomes.
The sandbox is not broken. It is simply being asked to do too much. By understanding its structural constraints—the data deficit, the institutional silos, the fixed-cost nature of regulation—we can design better systems around it. The goal is not to abandon sandboxes but to embed them in a more coherent strategy for regulatory adaptation. That strategy must include clear pathways to full authorization, proportional rulemaking, and a commitment to learning that extends well beyond the sandbox walls.
Recent Comments