The commitment without a threshold: Falsifiability conditions in sovereign-level performance agreements
Keywords: Performance Measurement; Falsifiability; Sovereign Commitments; Natural Process Limits; Development Finance; Bilateral Agreements; Signal Architecture
Preferred citation: Skogvold, T. (2026). The commitment without a threshold: Falsifiability conditions in sovereign-level performance agreements. MetriqOne Publishing.
Abstract
Bilateral agreements and multilateral development commitments routinely produce performance promises — railway networks, border infrastructure, economic cooperation zones, technology partnerships — without attaching a measurement architecture capable of distinguishing a process on track for delivery from one that has quietly drifted off it. This paper argues that the consistent absence of formalised Natural Process Limits in sovereign-level performance agreements leaves committed outcomes without a basis for falsifiable evaluation, and leaves genuine delivery indistinguishable from non-delivery until the window for intervention has closed. Drawing on Wheeler and Chambers’ (1992) signal-versus-noise methodology, Popper’s (1959) falsifiability standard, and Deming’s (1986) work on the distinction between systemic and individual causes of variation, this paper proposes that a pre-signal threshold — set voluntarily by the committing party at the moment of agreement — gives a government the earliest and clearest available means of demonstrating, in verifiable terms, that its commitment is being delivered. A 2026 bilateral infrastructure communiqué between two Asian governments is examined as an illustrative — not conclusive — case in point. The paper is addressed to the signatories themselves: the development bankers, bilateral agreement architects, and institutional fund officers who have not, to the author’s knowledge, been offered this architecture before, and who stand to gain the most from adopting it first.
1. Introduction
There is a category of document that looks like a performance commitment but functions as something else entirely. It names outcomes. It specifies cooperation areas. It records mutual intent. It is signed at a high level, announced at a press conference, and filed in the institutional record of two or more governments.
What it does not contain is a threshold.
By threshold, this paper means something precise: a pre-agreed, formally recorded level of observable evidence below which the committing party has determined, in advance, that the commitment is failing. Not a target. Not an aspiration. A falsifiable boundary — the structural condition that separates a scientific claim from a political declaration.
Without that threshold, the commitment cannot be falsified. And a commitment that cannot be falsified, in Popper’s (1959) terms, is not a claim about the world. It is a statement about intent — which is a different category of thing entirely.
This paper does not argue that the signatories of such agreements are acting in bad faith. The stronger argument is the opposite: most signatories are acting in good faith, with the institutional tools they were given. The tools do not include a falsifiability condition. Nobody told them one existed. The result is that a genuine commitment and a stalled one produce, for years, exactly the same observable record — an announcement, a timeline, and silence — until it is far too late to tell them apart.
That is the problem this paper addresses.
2. The Architecture of a Sovereign-Level Performance Commitment
A bilateral agreement, a multilateral development framework, or an infrastructure cooperation communiqué shares a structural anatomy. It contains a statement of intent, a list of cooperation areas, a set of mechanisms for operationalising the cooperation, and — in the more sophisticated examples — a timeline and a designated institutional vehicle for delivery.
What it does not contain, as a matter of consistent structural practice, is any of the following:
- A baseline reading against which progress can be measured.
- A process limit — upper or lower — that defines the boundary between normal variation and a signal worth acting on.
- A pre-signal threshold set at or below the lower process limit, the breach of which triggers an early warning before the process has failed.
- A falsifiability condition: a pre-agreed statement of what evidence, at what level, would constitute confirmation that the commitment has not been met.
The absence of these four structural elements is not accidental. It reflects the institutional inheritance of performance management frameworks that themselves lack these conditions — the Balanced Scorecard, OKRs, and LogFrame among them (Neely, Gregory and Platts, 1995; Franco-Santos et al., 2007). Behn (2003) made a related point about public-sector performance measurement specifically: different management purposes — improving, controlling, budgeting, motivating, evaluating — require different measures, and conflating them produces systems that serve none of these purposes well. A bilateral agreement asks a single set of stated commitments to serve evaluation, motivation, and diplomatic signalling simultaneously, with no measurement architecture differentiated enough to do any of the three with precision. If the frameworks that governments and development institutions use to manage their own internal performance do not embed Natural Process Limits, it follows that the commitments they make to each other will not embed them either.
The problem is structural, not personal. It originates in the design of the tools, not in the character of the people using them.
3. The Seventy-Year Gap at Sovereign Scale
Since Drucker’s (1954) foundational proposition that organisations require a balanced set of performance measures, the field of performance measurement has expanded across disciplines without achieving the architectural coherence required to produce falsifiable evidence. Neely et al. (1995) documented the fragmentation. Franco-Santos et al. (2007) confirmed the absence of a unified definition. Bititci et al. (1997) identified the alignment gap: strategic intent does not reliably reach field-level activity.
Deming (1986) and Wheeler and Chambers (1992) had already solved the signal-versus-noise problem in engineering and quality management, building on Shewhart’s (1931) original statistical distinction between common-cause and special-cause variation. Their solution — the XmR chart with its Natural Process Limits — provides a mechanism for distinguishing a process in statistical control from a process producing a genuine signal. But their work did not migrate into the performance measurement frameworks that governments and development institutions adopted. The field lacked a falsifiability standard (Popper, 1959), and it built its institutional infrastructure without one.
Lakatos’ (1970) refinement of Popper’s standard is useful here. Lakatos observed that a single disconfirming data point rarely overturns a research programme in practice; what matters is whether the programme’s protective belt of auxiliary assumptions can keep absorbing anomalies without degenerating into ad hoc rescue. A sovereign performance commitment without a threshold cannot even reach this test. There is no auxiliary assumption to defend and no anomaly to absorb, because there was never a prediction precise enough to be embarrassed by the evidence. It fails a weaker standard than the one Lakatos was worried about.
At the organisational level, this gap produces measurement systems that cannot distinguish between noise and a signal worth acting on. At the sovereign level, it produces something more consequential: performance commitments that cannot be evaluated at all.
When a government commits to building a railway, establishing a cross-border economic zone, or deploying smart border infrastructure, the commitment enters the record without a baseline, without process limits, and without a pre-signal threshold. There is no mechanism to determine, at any point before completion or abandonment, whether the commitment remains on a normal delivery trajectory or has already drifted off it.
This is Falsifiability Condition Four (FC4) operating at sovereign scale: the structural absence of formalised Natural Process Limits in a performance commitment framework (Skogvold, 2026). In an SME or an NGO, FC4 produces management decisions based on noise rather than signal. In a bilateral agreement governing billions in infrastructure investment and decades of sovereign obligation, it produces something with wider consequences.
4. The Two Records Problem
There is a distinction that performance measurement frameworks have never been required to operationalise, because they have never had the architectural tools to do so. It is the distinction between a commitment that is genuinely on track and one that has quietly stalled.
A commitment can encounter conditions — technical, financial, political, environmental — that prevent its realisation on schedule. This is a process outcome, explainable after the fact by reference to a trajectory that departed from what was intended. A separate commitment can stall for entirely different reasons, never reaching the funding, institutional capacity, or sustained political attention required to move at all. From the outside, in the absence of continuous evidence, these two situations produce an identical public record: an announcement, a signed agreement, a timeline, and then years of silence.
The critical observation is this: without a pre-signal threshold, a delayed-but-genuine commitment and a stalled one are indistinguishable from the public record alone, for as long as both parties choose not to say more.
A pre-signal threshold — set voluntarily, at the point of commitment, by the committing party — changes this. It is a formal declaration, made in advance, of the observable level below which the committing party has determined that the commitment is not proceeding as intended. Its primary function is not detection of wrongdoing. It is demonstration of delivery. A government that sets a threshold and stays inside it has produced something a communiqué alone cannot: verifiable, contemporaneous evidence that the commitment is being honoured, available long before completion and immune to retrospective dispute.
This is, in essence, a commitment device in Schelling’s (1960) sense: a constraint a party voluntarily accepts in advance precisely because it makes a stated intention more credible than the intention alone could be. The international relations literature on audience costs (Fearon, 1994) makes a related argument about why public commitments carry weight at all — a government that stakes its credibility on a public statement has more to lose from visible non-delivery than one that has said nothing. But audience costs alone are a blunt instrument: they only bite once non-delivery is already visible to the audience, which in infrastructure terms can be a decade after the commitment was made. A pre-signal threshold sharpens the same mechanism considerably, attaching the cost of non-delivery to an early, verifiable reading rather than to the distant and ambiguous moment when a public finally notices a railway was never built.
Set against a counterpart who has not adopted a threshold, the party that has done so holds the stronger position — not because it has exposed the other, but because it has produced evidence where the other has produced only an announcement. The threshold is, in this sense, a competitive instrument as much as a diagnostic one: the fastest available route to demonstrating that one’s own commitment was real.
5. The Giga Project Pattern
The literature on large-scale infrastructure projects in development contexts documents a remarkably consistent pattern. Flyvbjerg et al. (2002) demonstrated that cost overruns in major infrastructure projects are the rule, not the exception — systematic, not random, and consistently underestimated at the point of commitment. A later survey of the evidence (Flyvbjerg, 2014) found the pattern holding across decades and continents, regardless of the financing model or the political system involved — which suggests the gap this paper identifies is general, not specific to any one type of lender, borrower, or region. Hirschman (1967) identified what he called the Hiding Hand: the tendency of project appraisals to conceal the true difficulty of a project behind optimistic projections, on the grounds that decision-makers would not commit to projects if the full complexity were visible at the outset. Easterly’s (2006) broader critique of development aid effectiveness reaches a parallel conclusion from a different direction: large, top-down commitments routinely outrun the institutional capacity available to deliver them, and the resulting gap between announcement and outcome is absorbed quietly rather than corrected.
The sovereign financing dimension of this pattern has its own literature, separate from project management. Reinhart and Rogoff (2009) traced centuries of sovereign borrowing cycles in which financing commitments made in confident terms were serviced, at the margin, by the borrowing government long after the originating political moment had passed. More recent and more directly relevant to bilateral infrastructure lending, Brautigam’s (2009) empirically grounded account of Chinese development financing in Africa found the popular debt-trap narrative considerably oversimplified — financing terms, renegotiation patterns, and outcomes varied widely by project and by country, in ways a single narrative of predatory lending does not capture. That finding cuts in this paper’s favour, not against it: it indicates that outcomes in this category of agreement are genuinely uncertain at the point of signing, for reasons that have nothing to do with either party’s good faith — which is precisely the condition a falsifiable, threshold-equipped commitment is designed to resolve, by replacing both the optimistic narrative and the suspicious one with a verifiable record.
The pattern across development-context giga projects runs as follows. Leaders meet and commit to transformative infrastructure — railway connectivity, cross-border economic corridors, smart logistics systems, energy interconnection. The commitment is framed as a concrete action plan. Framework agreements are signed. Feasibility studies are commissioned. Contractors are selected. Ground is broken at a ceremony with significant press coverage. The project then encounters delays, cost revisions, and scope reductions. The original timeline is quietly extended. Political attention moves elsewhere. The project is eventually completed in reduced form, delivered on a longer timeline than announced, or scaled back from its original ambition.
In each case, the sovereign obligation taken on at signing is real, regardless of the eventual delivery outcome. But the measurement architecture that would have made the departure from the committed trajectory visible early enough to act on — by either party — was never attached to the agreement that created the obligation.
This is not a claim that any particular party intends a different outcome than the one announced. It is a claim about what is missing from the agreement: a mechanism by which either government could demonstrate, while the project is underway, that delivery remains on the committed path. A threshold-equipped commitment would not prevent every difficulty a major infrastructure project encounters. But it would give both signatories — and their domestic publics — a continuous, verifiable basis for confidence in the commitment, available well before the ribbon-cutting and the silence that sometimes follows it.
6. Case Exhibit: A 2026 Bilateral Infrastructure Communiqué
In June 2026, a phone call between the leaders of two Asian governments, conducted shortly after a domestic legislative transition in one of the two countries, produced a communiqué that is structurally representative of the category this paper examines.
The commitments recorded include: acceleration of railway cooperation and multimodal transport connectivity; deployment of smart border gate models; establishment of cross-border economic cooperation zones; expansion of high-tech investment; development of energy connectivity; a multi-year tourism cooperation initiative; and coordination during upcoming regional summit cycles.
Each commitment is a performance claim. Each names an outcome. Each implies a timeline, a responsible party, and a delivery standard. As recorded, none attaches a baseline. None specifies a process limit. None records a pre-signal threshold.
This is presented as an illustrative example of a structural pattern documented in the wider literature (Flyvbjerg, Holm and Buhl, 2002; Hirschman, 1967), not as a conclusive case against either government. The communiqué is a public diplomatic document, not an operational delivery plan, and the absence of a measurement architecture from its text does not establish that no such architecture exists in the implementing agencies’ internal planning. What the example illustrates is an opportunity rather than a deficiency: whichever of the two governments first attaches a baseline, process limits, and a pre-signal threshold to its share of the commitment gains a visible, verifiable basis for demonstrating that its delivery is on track — a stronger position, in evidentiary terms, than the announcement alone can provide.
The railway commitment is the most structurally significant element. Rail connectivity of this kind involves financing architecture, contractor selection, land acquisition, environmental assessment, and operational economics — each a delivery dependency in its own right. Attaching a falsifiability condition to this specific commitment — a pre-agreed statement of what evidence, at what level, would confirm the commitment is on track — would let either or both governments demonstrate delivery in terms that a communiqué cannot.
This observation does not depend on, and does not make, any claim about either government’s intentions. It is a structural observation about the tools both governments were given by the field of performance measurement as it currently stands — tools that have not, to date, embedded Natural Process Limits as a standard feature of sovereign-level agreements anywhere in the world. The opportunity to be the first to close that gap is open to either party, or both.
7. What a Threshold-Equipped Agreement Looks Like
A threshold-equipped bilateral performance agreement differs from a conventional communiqué in four structural ways.
First, it establishes a baseline. For each committed outcome, the baseline records the observable starting condition — the current state of the railway network, the current processing capacity at the target border crossing, the current volume of bilateral trade in the specified sectors. The baseline is not aspirational. It is a reading.
Second, it establishes process limits. Drawing on the statistical methodology of Wheeler and Chambers (1992), upper and lower Natural Process Limits define the range within which normal variation in delivery progress is expected to occur. A reading within these limits is noise. A reading outside them is a signal — evidence that the delivery process has departed from its expected trajectory.
Third, it establishes a pre-signal threshold. Set at or below the lower Natural Process Limit, the pre-signal threshold is the earliest warning layer. It is set voluntarily by the committing party at the point of agreement — not imposed externally, not calculated after delivery has begun. Its function is to give the committing party maximum lead time: weeks or months of warning before the statistical limit is breached, before the signal becomes unambiguous, before the intervention window has closed.
Fourth, it establishes a falsifiability condition. The committing party specifies, in advance, what evidence at what level would constitute confirmation that the commitment has not been met. This is the scientific floor of the agreement — the condition that elevates it from a political declaration to a claim about the world that can be confirmed, qualified, or denied by reference to observable data.
None of these four structural conditions require new institutional machinery at the point of enforcement. They require an agreement that records them at the point of commitment. The architecture is not complex. The resistance to it is not technical.
8. The Deal Worth Signing
This paper is addressed to the signatories. Not as a critique of what they have signed, but as an account of what they were not given.
The development banker who structures a bilateral infrastructure facility did not design the institutional framework in which sovereign commitments operate without falsifiability conditions. The bilateral agreement signatory who commits to railway connectivity did not choose to exclude a pre-signal threshold from the agreement architecture. The multilateral fund officer who approves a cross-border economic zone did not decide that the performance conditions attached to the disbursement would not include Natural Process Limits.
They were given tools that do not include these conditions. And in the absence of a threshold, when the project does not arrive — when the railway is not built, or is built at a fraction of its committed capacity, or is built and operated at a loss that the host government services for thirty years — there is no contemporaneous record showing when the trajectory first departed from the commitment, or what could have been done about it at the time.
The threshold changes this. Not for the project. For the signatory.
A signatory who insists on a pre-signal threshold as a condition of agreement is a signatory who can prove their delivery, in real time, rather than asking their counterpart and their own public to wait for a ribbon-cutting that may or may not arrive on schedule. The threshold is not a mechanism of blame. It is the fastest available route to verifiable evidence — the difference between announcing intent and demonstrating performance.
The threshold-equipped agreement is, in evidentiary terms, the stronger document. For the signatory who is confident in their own delivery, it is close to costless to adopt and considerably more persuasive than the announcement alone. For a counterpart who has not adopted one, the contrast does the work that no accusation needs to.
That is the deal worth signing. The field of performance measurement has not, to the author’s knowledge, offered this architecture to sovereign-level signatories in this form before.
MetriqOne is proposed as a framework that satisfies the structural conditions this paper describes — embedding falsifiability, Natural Process Limits, and the pre-signal threshold as design conditions rather than optional additions (Skogvold, 2026). The contribution of this paper is not to advocate for a particular product. It is to establish, as a matter of academic record, that the structural conditions for a falsifiable sovereign-level performance commitment exist — and that their consistent absence from bilateral agreements, development finance instruments, and multilateral cooperation frameworks is a gap in the field, available to be closed by whichever institution moves first.
9. Conclusion
The performance commitment without a threshold is the dominant form of sovereign-level agreement. It is not dominant because it is the best available option. It is dominant because the field of performance measurement has not provided an alternative.
This paper has argued that the consistent structural absence of Natural Process Limits and pre-signal thresholds in bilateral and multilateral performance commitments leaves committed outcomes without a basis for falsifiable evaluation, and leaves genuine delivery and stalled delivery operationally indistinguishable until it is too late to act on the difference.
The argument has three implications. For scholars: the seventy-year gap in performance measurement’s disciplinary architecture has consequences at the sovereign level that have not been examined in the literature. The extension of FC4 — the absence of formalised Natural Process Limits — from the organisational to the sovereign scale is a research agenda, not a footnote.
For institutions: the bilateral agreement, the development finance instrument, and the multilateral cooperation framework are performance documents. They can be strengthened, at negligible cost, by the structural conditions of falsifiable performance claims — without requiring any change to the diplomatic language in which they are written.
For signatories: the threshold is not a threat. It is the architecture that turns a declaration into demonstrable evidence — and the surest, lowest-cost way for a government confident in its own delivery to show it.
References
Behn, R.D. (2003) ‘Why measure performance? Different purposes require different measures’, Public Administration Review, 63(5), pp. 586–606.
Bititci, U.S., Carrie, A.S. and McDevitt, L. (1997) ‘Integrated performance measurement systems: A development guide’, International Journal of Operations and Production Management, 17(5), pp. 522–534.
Brautigam, D. (2009) The Dragon’s Gift: The Real Story of China in Africa. Oxford: Oxford University Press.
Deming, W.E. (1986) Out of the Crisis. Cambridge, MA: MIT Press.
Drucker, P.F. (1954) The Practice of Management. New York: Harper & Row.
Easterly, W. (2006) The White Man’s Burden: Why the West’s Efforts to Aid the Rest Have Done So Much Ill and So Little Good. New York: Penguin Press.
Fearon, J.D. (1994) ‘Domestic political audiences and the escalation of international disputes’, American Political Science Review, 88(3), pp. 577–592.
Flyvbjerg, B. (2014) ‘What you should know about megaprojects and why: An overview’, Project Management Journal, 45(2), pp. 6–19.
Flyvbjerg, B., Holm, M.S. and Buhl, S. (2002) ‘Underestimating costs in public works projects: Error or lie?’, Journal of the American Planning Association, 68(3), pp. 279–295.
Franco-Santos, M. et al. (2007) ‘Towards a definition of a business performance measurement system’, International Journal of Operations and Production Management, 27(8), pp. 784–801.
Hirschman, A.O. (1967) Development Projects Observed. Washington, DC: Brookings Institution.
Lakatos, I. (1970) ‘Falsification and the methodology of scientific research programmes’, in Lakatos, I. and Musgrave, A. (eds.) Criticism and the Growth of Knowledge. Cambridge: Cambridge University Press, pp. 91–196.
Neely, A.D., Gregory, M.J. and Platts, K.W. (1995) ‘Performance measurement system design: A literature review and research agenda’, International Journal of Operations and Production Management, 15(4), pp. 80–116.
Popper, K.R. (1959) The Logic of Scientific Discovery. London: Hutchinson.
Reinhart, C.M. and Rogoff, K.S. (2009) This Time Is Different: Eight Centuries of Financial Folly. Princeton: Princeton University Press.
Schelling, T.C. (1960) The Strategy of Conflict. Cambridge, MA: Harvard University Press.
Shewhart, W.A. (1931) Economic Control of Quality of Manufactured Product. New York: Van Nostrand.
Skogvold, T. (2015) Performance Measurement in Small and Medium Enterprises: A Study of Measurement Frameworks in Volatile Environments. MSc Dissertation. University of Liverpool.
Skogvold, T. (2026) MetriqOne: Mental software and the emergence of a new discipline in performance measurement. MetriqOne Publishing.
Wheeler, D.J. and Chambers, D.S. (1992) Understanding Statistical Process Control. 2nd edn. Knoxville, TN: SPC Press.
Wheeler, D.J. (2000) Understanding Variation: The Key to Managing Chaos. 2nd edn. Knoxville, TN: SPC Press.
The MetriqOne Trilogy — the full argument in three volumes.
Get the Trilogy on Amazon →