What would real evaluation of EPR performance require?¶
The four preceding articles assemble an uncomfortable composite: no causal study of EPR's recycling effect anywhere; measurement systems that can move reported performance by thirty points without touching a bale; a comparative record whose league tables dissolve into basis differences; and an environmental-release promise with no supporting study for the instrument's dominant form. This closing article treats that composite as a design problem rather than a complaint. It asks why a policy field moving billions annually runs on unevaluated claims; documents how the current systems destroy their own evaluability — sometimes carelessly, sometimes conveniently; specifies what credible evaluation would require, none of it methodologically novel; inventories the natural experiments now running unused; and proposes the minimal evaluation agenda that would, within five years, replace assertion with knowledge on the field's central questions.
1. Why the field runs on unevaluated claims¶
The evaluation gap is not an accident of youth — packaging EPR is thirty-five years old — and the incentive analysis explains its persistence better than any technical difficulty.
Nobody with the data has a client-free interest in the answer. Producer organisations hold declaration-level supply data, collection records and fee models; their interest is in demonstrating success and containing scrutiny. Regulators hold reported data; their capacity is documented in Theme 2, and an evaluation that found their programs wanting would indict their oversight. Industry funders of counter-research want cost findings, not performance findings; advocacy funders want the reverse. The academic community lacks data access: the declaration-level records that would support producer-level analysis sit behind the confidentiality walls documented in the transparency article. The result is the pattern visible across Themes 4–6: advocacy-adjacent modelling on all sides, and a research base whose best causal studies (the fee-rate null; the pass-through analogues) were possible only where data escaped the walls.
The claims serve their purposes unevaluated. "EPR raises recycling rates" recruits legislators; "EPR is a hidden grocery tax" recruits opponents; neither constituency needs the claim tested, and both are served by its untestability. The unfalsifiability documented for free-riding (Theme 4) generalises to the performance domain.
2. How the systems destroy their own evaluability¶
The record contains a disquieting pattern: at the moments when evaluation would have been most possible, the measurement apparatus was dismantled or the basis changed.
Ontario is the fullest case, assembled across this library: the Continuous Improvement Fund and municipal cost Datacall — the instruments that produced North America's longest consistent cost-and-performance series — terminated in July 2025, at the transition's completion; the producer-audit requirement waived for the foundational 2025–26 data years; the new system reporting on a different basis (steward-reported supply) than the old (estimated generation), breaking the series at exactly the point a before/after comparison would bite; and the post-transition rate still unpublished. The largest EPR transition in North American history — the near-ideal natural experiment, with a hard switch date and adjacent-province controls — has been rendered substantially unevaluable by choices made during it.
The European methodology break (2020) was, by contrast, an evaluability improvement purchased at the cost of the time series: the stricter calculation point produced better numbers and an eleven-year series that cannot be read across the break. Only four member states maintained dual reporting — the cheap continuity measure that would have preserved both.
The UK enters its modulated-fee era with the same structural choice made worse: an eleven-point gap between its two 2024 methodologies, anticipatory price movements smearing the pass-through event window, and no committed evaluation of the flat-to-modulated fee switch — the cleanest modulation experiment ever run, launching unobserved.
The routine reporting layer compounds the transition failures. Even where measurement instruments survive, their outputs arrive too late and too unreliably to support evaluation: Ontario's regulator was publishing performance data roughly twelve months behind its own three-month target, with 24% of the collection sites on its public map found not to accept the materials listed for them; French scheme data, per the state inspectorates, runs about two years behind and incomplete; and Ontario's 2024 collection data was still pending publication in mid-2026, six months after the transition it would describe. An evaluation community cannot form around data that arrives years late, changes basis, and cannot be reconciled to what preceded it — which is part of why the community has not formed.
The generalisation: evaluability is a casualty of transitions unless explicitly protected, and no jurisdiction has yet protected it. Whether by budgetary neglect (Ontario's data instruments), reform zeal (the EU's break), or the quiet convenience of unmeasurable baselines, the systems keep arriving at their decisive moments without the instruments that would judge them.
3. What credible evaluation requires¶
None of the requirements is novel; all are standard in program evaluation, and their absence from EPR is the finding.
Stable, dual-reported metrics across any transition. When a basis must change, report both bases for the overlap years — the EU's four-state practice, generalised. Cost: trivial. This single rule would have preserved the Ontario and European series.
Baselines established before program start, on the program's future basis. Oregon launched with a state recovery rate on one basis and a program that will report on another; British Columbia has no pre-2014 baseline at all. A jurisdiction six months from launch can still fix this; none has.
Independent access to declaration-level data under confidentiality. Producer-level supply and fee-exposure data, accessible to qualified researchers under the three-tier disclosure structure of Theme 2, would enable the producer-level design-response analysis that thirty-five years of modulation debate has never had (Theme 5).
Pre-registered evaluation attached to every reform. The Netherlands doubling a differential, Quebec escalating a malus, the UK switching to modulated fees — each is an experiment whose evaluation design should have been filed before the switch, with the difference-in-differences comparators named.
Audit of the measurement chain itself. Contamination sampling, yield tracking through sorting and reprocessing, and export verification — the recycled-at-scale apparatus the PPWR now mandates from 2035, which is best understood as the EU forcing the measurement infrastructure its evaluation always lacked.
4. The natural-experiment inventory, running unused¶
The field's standing irony is that its evaluation drought coincides with the richest experimental period in its history. The inventory, as of 2026:
| Experiment | Design opportunity | Status |
|---|---|---|
| Ontario transition (completed Jan 2026) | Hard switch, large market, adjacent-province controls — the flagship | Compromised (Section 2); partially recoverable if the new series is published with overlap reconstruction |
| UK flat→modulated fees (2025–26 → 2026–27) | The cleanest modulation dose test ever run | No committed evaluation |
| Oregon launch (July 2025) | First US program; scanner-data pass-through study feasible | First-year report silent on prices and rates |
| Quebec 2027 malus escalation (20%→75%/50%) | The strongest dose experiment anywhere | No announced evaluation |
| Netherlands doubled recyclate discount (2025) | Producer-level uptake response measurable in scheme data | Scheme-internal; unpublished |
| Ireland DRS launch (Feb 2024) | Litter and return-rate effects — already yielding results | The exception: national litter surveys are capturing it |
| EU export ban (Nov 2026) | Stress test of export-dependent recycling rates | Effects will surface in 2027–28 statistics, evaluated or not |
| PPWR grades (2028–2030) | Continental modulation-with-bans experiment | Evaluation architecture unbuilt; delegated acts delayed |
Ireland demonstrates what the others lack: an independent measurement institution (the national litter survey) that predates the intervention, persists through it, and publishes on its own schedule. The lesson is institutional, not methodological.
5. The minimal agenda¶
Five studies, ordered by feasibility, that would convert the field's central assertions into findings within five years:
- The Ontario event study — scanner-price and municipal-budget analysis across the January 2026 switch, with Manitoba and the Atlantic provinces as controls; recoverable now, degrading with every year of anticipatory adjustment.
- The UK modulation evaluation — producer-level fee-exposure against packaging-portfolio change across the Year 1/Year 2 switch, using the administrator's own data; the state holds everything required.
- A net-incidence accounting for one jurisdiction — the Theme 4 specification: pass-through estimate, municipal tracing, decile weighting; Washington's visible-billing structure makes it the natural site.
- A misreporting estimate from any audit-active scheme — France's 15% audit stream, published as a sampled correction rate; the number exists inside Citeo today.
- The dual-basis reconstruction of one major series — re-stating either the Ontario or the German series across its break, establishing the method for every future transition.
The estimated cost of all five is a rounding error against the systems' annual budgets — the UK scheme's administration alone runs £28 million a year. The barrier is not resources but the incentive structure of Section 1, which is why the agenda's real precondition is institutional: evaluation must be commissioned by parties outside the claim structure, with data access mandated rather than negotiated.
The record already identifies who those parties are, because nearly every independent finding in this library came from the same institutional family. The state audit offices produced the field's best evidence — the Auditor General of Ontario's enforcement audit, the French inspectorates' governance findings, the UK National Audit Office's export-evidence analysis, the European Court of Auditors' target assessment — and they hold two properties no other evaluator combines: statutory data access that schemes cannot refuse, and no stake in the programs' success. Statistical agencies are the natural custodians of the dual-reporting and series-continuity rules — Eurostat's methodology governance, for all the disruption of its 2020 break, is the only apparatus in the field that documents its own measurement choices. And independent survey institutions demonstrate the third role: Ireland's litter surveys captured the deposit launch's effects because an instrument with no connection to the scheme happened to predate it. A jurisdiction serious about evaluation does not need to invent institutions; it needs to task these three — audit office, statistical agency, standing survey — before the next reform, with the evaluation designs pre-registered and the scheme's data access written into the program's legal instrument rather than left to its goodwill.
One further design is worth naming because it resolves the confidentiality objection that schemes will otherwise raise: the research data trust — declaration-level data deposited under the regulator's confidentiality tier (Theme 2's three-tier structure), accessible to accredited researchers under the disclosure-control rules national statistical agencies have operated for decades for tax and health microdata. The precedents are mature; the packaging application is merely unbuilt.
6. Where the argument stands — and a closing word for the reader¶
This theme set out to answer whether EPR works, and has ended somewhere more precise: the field's performance questions are answerable, and the field is organised not to answer them. The methods exist; the experiments are running; the data sits in named institutions; and the cost of knowing is trivial. What stands between assertion and knowledge is a structure in which every party holding data has a stake in the claims, transitions consume the measurement instruments that would judge them, and the reported numbers — collected, counted and denominated by the systems under evaluation — are built to flatter.
For the reader, the practical residue of this theme is a posture rather than a verdict. Treat every EPR performance claim as carrying an implicit measurement appendix, and demand it: the basis, the break points, the counting location, the denominator's source, the export share, the evaluator's client. Where the appendix is refused, the claim is advertising. Where it is supplied, the field's genuine achievements — funded collection at scale, the German plastics-quota trajectory, the deposit litter record, the EU's accounting reform — survive scrutiny comfortably, which is itself the demonstration that honesty and the instrument's case can coexist. The systems that will deserve confidence in 2030 are the ones now writing their evaluation designs before their reforms — and as of this writing, that set is empty. It is the cheapest important thing any jurisdiction in this field could change tomorrow.
References¶
- The findings synthesised here are documented, with sources, in the four preceding articles of this theme and in Themes 2, 4 and 5; primary citations are held there rather than repeated.
- On the Ontario data-instrument terminations and audit waivers: RPRA notices (July 2025; November 2024) and the Auditor General of Ontario (December 2025), as documented in Themes 2 and 6.
- On dual reporting: Eurostat env_waspac metadata (the four-member-state dual-approach note).
- On the natural-experiment inventory: the scheme and statutory documentation cited in Themes 5 and 6; Lakhan, C. (2026), Ecomodulation of Extended Producer Responsibility Fees: A Literature Review — the unused-experiment inventory this article extends.
- On evaluation institutions: the Irish national litter-survey series (IBAL); the audit-office exemplars (Auditor General of Ontario; IGF/IGEDD/CGE; the UK National Audit Office; the European Court of Auditors) whose work supplies most of the independent findings in this library.
Verification note: this article is a synthesis; every empirical claim traces to the sourced articles it cites, and its recommendations are identified as this library's analysis rather than as findings. See Sources and method.