The thesis
Evidence → interpretation → proposal
A working prototype is evidence of something. The difficulty is deciding what. It may demonstrate implementation skill, effective collaboration, a well-chosen toolchain, or an unusually persuasive presentation. It does not, by itself, establish that the team found the right problem, improved the surrounding system, or understood the conditions under which its intervention would fail.
This is not an argument that hackathons have become pointless. Research describes a family of formats serving learning, community, innovation and other purposes. A format designed to bring people together need not behave like a professional examination. The error begins when one kind of evidence is used to support a stronger claim than the event can justify. [1]
What changed—and what has not been established
AI can accelerate bounded implementation tasks. A controlled Copilot study found a 55.8% reduction in completion time for its programming task. That is not a universal multiplier for engineering. METR’s early-2025 experiment found experienced developers took 19% longer on work in familiar repositories. Its 2026 update reports that selection effects make newer estimates unreliable. These findings describe different settings; neither establishes the value of every current tool. [2], [3], [4]
Our interpretation: implementation speed is becoming a less stable basis for comparison. The relevant question is not whether building has ceased to matter. It is whether a challenge can distinguish producing an artifact from making a justified improvement, while allowing participants to use the tools they actually have.
The environment is not the invention
Cyber ranges, autonomous cyber challenges, robotics competitions and simulation already evaluate behaviour inside systems. Chaos engineering already tests hypotheses through disruption. Engineering assessment already values problem formulation and experimentation. MERENIC cannot claim these ideas as discoveries. [5], [7], [9], [13], [15]
Our proposed contribution is an assessment contract that follows a claim across its life: what was observable when it was made; which alternative was rejected; what consequence was predicted; what changed after deployment; and whether the finding survived a controlled challenge and independent replay. The contract is portable. The environment can be simple.
A claim should carry its conditions
Consider a simulated dispatch network. Lower average response time could conceal worse service in an outlying district. An apparent improvement could also come from easier demand, a changed baseline or selective reporting. A convincing demonstration must make these rival explanations inspectable.
MERENIC therefore proposes three linked records. A decision receipt commits a prediction, alternatives and a stopping rule before the relevant test. A paired consequence record compares intervention and baseline under matched external conditions, including harms and costs. A revision trace connects a later decision to the evidence that changed it—or justifies retaining it. These are proposed protocol components, built on prior methods such as preregistration and controlled comparison. Their incremental value remains an empirical question. [21], [22]
Changing direction earns no automatic reward. A team could perform a theatrical failure and recovery. Equally, deciding not to deploy may be the strongest conclusion if the evidence does not support an intervention. What matters is the quality of the decision under the information and resources available at the time.
Measure consequences. Examine judgement.
Machines should check manifests, provenance, resource use and repeatability. The environment should emit observable consequences, with uncertainty and model limitations attached. Experts should examine whether the problem mattered, whether alternatives were investigated fairly, and whether the conclusion exceeds the evidence. A high aggregate score must never purchase permission to violate a declared safety constraint.
Reasoning records are evidence of declared decisions, not access to a person’s private cognition. AI can produce both excellent code and fluent explanations. We cannot promise to isolate an irreducibly human skill by reading those explanations. The primary assessment unit is the team with its declared tools. Where individual competence matters, it requires a separately validated assessment.
The standard must accept the same burden
Our hypothesis: this record of decisions and consequences will make technical evaluation more reproducible and harder to win through presentation alone. It may instead reward paperwork, domain familiarity, expensive infrastructure or teams skilled at anticipating the rubric. An honest first edition must be designed to detect those failures.
Edition 01 will therefore test the framework as well as its participants. Its protocol should be frozen before recruitment, its judgements independently scored, and its claims replayed under withheld conditions. The study must report disagreement, invalid runs, cost and adverse findings—not just successful interventions.
MERENIC is a research draft, not a validated certification. Its ambition is modest enough to be tested and serious enough to matter: when a team says it improved a system, make the evidence hold.
Prior-art landscape
This is a targeted, non-exhaustive prior-art review. The proposed gap is an interoperability and assessment-validity problem; absence of an identical format in this search is not proof of an unclaimed invention.
| Known practice | What exists | Limitation / unresolved question | Design response |
|---|---|---|---|
| Hackathons / corporate innovation | Multiple social, learning and innovation outcomes; prototypes and pitches are only part of the landscape. [1], [11] | A judgement of innovation is not automatically evidence of operational reliability. | Keep these purposes legitimate; limit MERENIC to evidence-bearing engineering claims. |
| CGC / AIxCC / attack–defence CTF | Deployed adversarial systems, discovery and repair, automated scoring. [5], [6] | Discovery and consequences are already measured in bounded security tasks. | Propose a reusable cross-domain record of participant-selected problems; do not claim the environment is new. |
| Cyber ranges / sandbox training | Instrumented simulated infrastructure and assessment. [7] | Environment fidelity and exercise success do not alone validate a portable judgement rubric. | Publish the claim-to-evidence mapping and replay contract. |
| Kaggle / measurable competitions | Metrics and held-out evaluation; adaptive leaderboard overfitting is researched. [8], [33] | Optimising a fixed objective can leave problem selection and external harms outside assessment. | Lock a selected outcome with an independent guardrail panel; disclose all evaluation slices. |
| RoboCup / rescue robotics | Agent behaviour assessed in rescue scenarios and simulation. [9] | This is a close precedent for the city example, not proof of novelty. | Evaluate the discovery and decision record separately from the controller outcome. |
| Digital twins / industrial simulation | Validation, uncertainty and interoperability are established concerns. [24], [25] | A sophisticated world can still be a poor model of the intended use. | Require a fidelity statement; no claim of real-world impact from simulation alone. |
| Living labs | Users co-create innovations in real-world settings. [10] | Ecological validity and repeatability pull in different directions. | Use synthetic conditions for comparison; later field validation requires separate protocols. |
| Assessment science / OSCE / ABET | Evidence-centred tasks, observed professional scenarios and engineering outcomes. [12], [13], [14] | Rich observation can still support an invalid inference about competence. | Specify constructs, train raters and test transfer before hiring or certification use. |
| Chaos / fault injection | Hypotheses about steady state tested through disruptive events. [15] | Breaking an implementation alone says little about selecting the right intervention. | Link stress families to declared assumptions, then run common held-out tests. |
| Reproducible research | Independent artifact evaluation and reproduction. [16] | Availability of code is weaker than reproduced results. | Require independent run identity, manifest hashes and tolerance checks. |
| AI-agent benchmarks | Repository tasks, fresh instances and benchmark mutation. [17], [18], [19] | Leakage, underspecified tests and benchmark-specific optimisation remain concerns. | Isolate scoring; use external observations, negative controls and fresh evaluation families. |
| AI-era hackathon redesign | HackerRank Orchestrate already centres AI-agent building and evaluation. [28] | Allowing AI and adding evaluation is not differentiation. | Treat human–AI attribution as unresolved and test protocol value directly. |
Mechanism ledger
Names below identify concrete records and procedures. None is asserted to be patentable or unprecedented. The candidate contribution is their explicit linkage and the evidence needed to validate that linkage.
| Mechanism | Prior art | Unresolved problem | Proposed mechanism | Test and novelty boundary |
|---|---|---|---|---|
| Decision receipt | Preregistration; evidence-centred assessment. [12], [21] | Final narratives can hide hindsight. | Timestamp a problem choice, alternatives, prediction, expected harms, budget and stopping rule before the intervention test. | Audit receipt order; compare blinded ratings with and without receipts. Not claimed novel in isolation. |
| Paired consequence record | Common random numbers and controlled comparisons. [22] | A changing world makes before/after comparisons misleading. | Run frozen baseline and intervention against an identical exogenous schedule. Keep intervention-dependent responses causal in each branch. | Report seed-wise deltas, uncertainty and guardrail vectors. Coupling may fail; disclose divergence. |
| Revision trace | Hypothesis testing and engineering decision records. [12], [21] | A polished account cannot show when contrary evidence became available. | Link each amended prediction to signed evidence and the next action, including justified non-action. | Blind raters judge warranted revision, not number of pivots. Incremental validity unproven. |
| Assumption challenge pairing | Chaos engineering, held-out tasks and evolving environments. [15], [19], [23] | Personalised stress can make participants incomparable. | Map declared dependencies to a public fault taxonomy; reserve a common holdout panel for ranked results. | Personalised faults are diagnostic only; measure transfer on the common panel. |
| Deferred-action test | Controlled experimentation; professional judgement assessment. [12], [14] | Build-only incentives penalise justified restraint. | Allow a no-change submission with evidence, predicted avoided harm and a falsifiable trigger for future action. | Baseline always runs. No automatic credit for abstention; review opportunity cost and rejected feasible choices. |
Research method
Research cut-off: 28 September 2026. Searches covered academic literature, official competition documentation, standards organisations, assessment bodies and technical documentation. Search families included hackathon outcomes, AI developer productivity, cyber-range assessment, adversarial scoring, scenario assessment, simulation validity, reproducibility, benchmark leakage, preregistration, environment co-evolution and standards governance.
Primary sources and research papers were preferred. Preprints are explicitly labelled. Company documentation establishes what a format claims to do; it does not independently establish validity. Search-result excerpts were sufficient for narrow landscape descriptions; linked documents support deeper inspection. This is not a systematic review, exhaustive patent search or independent replication.
Selection rule: retain sources that establish a close mechanism, challenge the central thesis, or change an implementation requirement. Reject promotional claims as evidence of superiority. An absence claim requires more than a missing search result. No empirical user data was collected for this draft.
Adversarial review & open questions
| Reviewer objection | Revision made | Still unresolved |
|---|---|---|
| CTO: “This is marketing for a simulator.” | Reduced the claim to a portable evidence contract; published all dependencies and the toy model’s limits. | Does the protocol change a real engineering decision? |
| Academic: “You have no construct validity.” | No certification, general ability score or human-only inference. Added blinded rating and transfer tests. | External validity, inter-rater reliability and domain bias. |
| Organiser: “This already exists.” | Named close precedents and dropped “first” and environment novelty claims. | Does the linked protocol add enough value to justify adoption? |
| Engineer: “I can game the metric.” | Added independent guardrails, matched baselines, negative controls, a frozen holdout and tamper boundaries. | Unknown proxies, simulator exploits and hidden distribution leakage. |
| AI researcher: “An agent can generate all your reasoning.” | Treat the team and its tools as the unit; receipts are commitments, not proof of cognition. | How to validate individual responsibility with increasingly autonomous agents? |
| Operator: “This is unaffordable.” | Capability profiles instead of prestige tiers; containers are sufficient for a narrow scenario. | Cost per valid replay, facilitator hours and repeated-run feasibility. |
| Participant: “The domain experts always win.” | Common briefing, practice environment, matched resources and separate domain reporting. | Accessibility, prior exposure and differential measurement effects. |
Additional failure modes: teams deliberately fail to earn adaptation credit; explore too many outcomes then report only one; exfiltrate hidden seeds; alter telemetry; use external paid agents beyond declared budgets; discover the simulator’s numerical shortcuts; or optimise the mean at the expense of minority outcomes. Report attempted gaming and rejected evidence, and allow an independent appeal.
References
Sources are not endorsements or affiliations. Dates and limitations below distinguish evidence from proposed methodology.
- On Hackathons: A Multidisciplinary Literature Review ↗CHI 2023 · peer-reviewed review; purposes, formats, processes and outcomes.
- The Impact of AI on Developer Productivity: Evidence from GitHub Copilot ↗Peng et al., 2023 · controlled task study; limited external generalisation.
- Early-2025 AI and experienced open-source developer productivity ↗METR, 2025 · randomised study in mature repositories.
- We are Changing our Developer Productivity Experiment Design ↗METR, February 2026 · selection effects limit newer estimates.
- Cyber Grand Challenge ↗DARPA · autonomous discovery, repair and deployment precedent.
- AI Cyber Challenge: final scoring guide announcement ↗DARPA, 2025 · functionality-preserving repair and automated scoring.
- The Cyber Range: A Guide ↗NIST-hosted guide · simulated environments and workforce assessment.
- Kaggle competition documentation ↗Kaggle · measurable submissions and competition evaluation.
- RoboCup Rescue ↗RoboCup Germany · simulation and embodied rescue assessment.
- Living Labs: origins, developments and future perspectives ↗ENoLL, 2025 · co-creation in real-world innovation settings.
- NASA Tournament Lab ↗NASA · institutional challenge-based problem solving.
- A Brief Introduction to Evidence-Centered Design ↗Mislevy, Almond & Lukas, ETS, 2003 · claims, evidence and assessment tasks.
- Engineering accreditation criteria 2025–2026 ↗ABET · problem formulation, experimentation and engineering judgement.
- A sample OSCE station ↗General Medical Council · scenario-based professional assessment.
- Principles of Chaos Engineering ↗Primary principles · steady-state hypotheses and realistic disruptive events.
- Lessons from Five Years of Artifact Evaluation at EuroSys ↗ACM SIGOPS, 2025 · independent artifact evaluation and reproduction.
- SWE-bench ↗Official benchmark · repository-level software issue resolution.
- SWE-rebench: decontaminated evaluation ↗2025 preprint · fresh tasks and contamination concerns.
- Saving SWE-Bench: A Benchmark Mutation Approach ↗2025 preprint · benchmark mutation already exists.
- Planning and Acting in Partially Observable Stochastic Domains ↗Kaelbling, Littman & Cassandra, 1998 · POMDP foundations.
- Preregistration ↗Center for Open Science · predictions distinguished from post-hoc explanation.
- Bayesian Optimization Allowing for Common Random Numbers ↗2019 preprint · shared stochastic conditions in comparison.
- LLM-POET: Evolving Complex Environments ↗2024 preprint · environment/agent co-evolution is prior art.
- Digital Twins for Advanced Manufacturing ↗NIST · validation, uncertainty and reference implementations.
- Simulation interoperability standards products ↗SISO · existing simulation interoperability standards.
- Guide to the IETF standards process ↗IETF · open review and running implementations; no affiliation implied.
- Conformity Assessment Basics ↗NIST · declaration, certification and accreditation are distinct.
- Behind the Scenes of HackerRank Orchestrate ↗HackerRank, 2026 · AI-agent hackathon and evolving evaluation.
- Multi-tenancy ↗Kubernetes · namespace isolation requires additional controls.
- OpenTelemetry signals ↗OpenTelemetry · metrics, logs and traces.
- Global Brand Database ↗WIPO · database coverage is incomplete; national searches also matter.
- Patent subject matter eligibility ↗USPTO · US eligibility guidance; not a patentability opinion.
- Reducing overfitting in challenge-based competitions ↗2016 preprint · adaptive leaderboard overfitting and Ladder precedent.