# MERENIC — research and design dossier

Research Draft 0.1 · 28 September 2026

MERENIC is a research initiative initiated by Majji Pradeep Kumar and developed with support from CyberMindSpace Private Limited. This is a targeted review, not an exhaustive literature review, trademark clearance or patent opinion. No novelty is asserted solely because this search did not find an identical protocol.


# Research

## The thesis


Evidence → interpretation → proposal

A working prototype is evidence of something. The difficulty is deciding what. It may demonstrate implementation skill, effective collaboration, a well-chosen toolchain, or an unusually persuasive presentation. It does not, by itself, establish that the team found the right problem, improved the surrounding system, or understood the conditions under which its intervention would fail.

This is not an argument that hackathons have become pointless. Research describes a family of formats serving learning, community, innovation and other purposes. A format designed to bring people together need not behave like a professional examination. The error begins when one kind of evidence is used to support a stronger claim than the event can justify. [[1]](research.html#ref-1)

### What changed—and what has not been established

AI can accelerate bounded implementation tasks. A controlled Copilot study found a 55.8% reduction in completion time for its programming task. That is not a universal multiplier for engineering. METR’s early-2025 experiment found experienced developers took 19% longer on work in familiar repositories. Its 2026 update reports that selection effects make newer estimates unreliable. These findings describe different settings; neither establishes the value of every current tool. [[2]](research.html#ref-2), [[3]](research.html#ref-3), [[4]](research.html#ref-4)

Our interpretation: implementation speed is becoming a less stable basis for comparison. The relevant question is not whether building has ceased to matter. It is whether a challenge can distinguish producing an artifact from making a justified improvement, while allowing participants to use the tools they actually have.

### The environment is not the invention

Cyber ranges, autonomous cyber challenges, robotics competitions and simulation already evaluate behaviour inside systems. Chaos engineering already tests hypotheses through disruption. Engineering assessment already values problem formulation and experimentation. MERENIC cannot claim these ideas as discoveries. [[5]](research.html#ref-5), [[7]](research.html#ref-7), [[9]](research.html#ref-9), [[13]](research.html#ref-13), [[15]](research.html#ref-15)

Our proposed contribution is an assessment contract that follows a claim across its life: what was observable when it was made; which alternative was rejected; what consequence was predicted; what changed after deployment; and whether the finding survived a controlled challenge and independent replay. The contract is portable. The environment can be simple.

### A claim should carry its conditions

Consider a simulated dispatch network. Lower average response time could conceal worse service in an outlying district. An apparent improvement could also come from easier demand, a changed baseline or selective reporting. A convincing demonstration must make these rival explanations inspectable.

MERENIC therefore proposes three linked records. A decision receipt commits a prediction, alternatives and a stopping rule before the relevant test. A paired consequence record compares intervention and baseline under matched external conditions, including harms and costs. A revision trace connects a later decision to the evidence that changed it—or justifies retaining it. These are proposed protocol components, built on prior methods such as preregistration and controlled comparison. Their incremental value remains an empirical question. [[21]](research.html#ref-21), [[22]](research.html#ref-22)

Changing direction earns no automatic reward. A team could perform a theatrical failure and recovery. Equally, deciding not to deploy may be the strongest conclusion if the evidence does not support an intervention. What matters is the quality of the decision under the information and resources available at the time.

### Measure consequences. Examine judgement.

Machines should check manifests, provenance, resource use and repeatability. The environment should emit observable consequences, with uncertainty and model limitations attached. Experts should examine whether the problem mattered, whether alternatives were investigated fairly, and whether the conclusion exceeds the evidence. A high aggregate score must never purchase permission to violate a declared safety constraint.

Reasoning records are evidence of declared decisions, not access to a person’s private cognition. AI can produce both excellent code and fluent explanations. We cannot promise to isolate an irreducibly human skill by reading those explanations. The primary assessment unit is the team with its declared tools. Where individual competence matters, it requires a separately validated assessment.

### The standard must accept the same burden

Our hypothesis: this record of decisions and consequences will make technical evaluation more reproducible and harder to win through presentation alone. It may instead reward paperwork, domain familiarity, expensive infrastructure or teams skilled at anticipating the rubric. An honest first edition must be designed to detect those failures.

Edition 01 will therefore test the framework as well as its participants. Its protocol should be frozen before recruitment, its judgements independently scored, and its claims replayed under withheld conditions. The study must report disagreement, invalid runs, cost and adverse findings—not just successful interventions.

MERENIC is a research draft, not a validated certification. Its ambition is modest enough to be tested and serious enough to matter: when a team says it improved a system, make the evidence hold.


## Prior-art landscape


This is a targeted, non-exhaustive prior-art review. The proposed gap is an interoperability and assessment-validity problem; absence of an identical format in this search is not proof of an unclaimed invention.

 | Known practice | What exists | Limitation / unresolved question | Design response

 | Hackathons / corporate innovation | Multiple social, learning and innovation outcomes; prototypes and pitches are only part of the landscape. [[1]](research.html#ref-1), [[11]](research.html#ref-11) | A judgement of innovation is not automatically evidence of operational reliability. | Keep these purposes legitimate; limit MERENIC to evidence-bearing engineering claims.

 | CGC / AIxCC / attack–defence CTF | Deployed adversarial systems, discovery and repair, automated scoring. [[5]](research.html#ref-5), [[6]](research.html#ref-6) | Discovery and consequences are already measured in bounded security tasks. | Propose a reusable cross-domain record of participant-selected problems; do not claim the environment is new.

 | Cyber ranges / sandbox training | Instrumented simulated infrastructure and assessment. [[7]](research.html#ref-7) | Environment fidelity and exercise success do not alone validate a portable judgement rubric. | Publish the claim-to-evidence mapping and replay contract.

 | Kaggle / measurable competitions | Metrics and held-out evaluation; adaptive leaderboard overfitting is researched. [[8]](research.html#ref-8), [[33]](research.html#ref-33) | Optimising a fixed objective can leave problem selection and external harms outside assessment. | Lock a selected outcome with an independent guardrail panel; disclose all evaluation slices.

 | RoboCup / rescue robotics | Agent behaviour assessed in rescue scenarios and simulation. [[9]](research.html#ref-9) | This is a close precedent for the city example, not proof of novelty. | Evaluate the discovery and decision record separately from the controller outcome.

 | Digital twins / industrial simulation | Validation, uncertainty and interoperability are established concerns. [[24]](research.html#ref-24), [[25]](research.html#ref-25) | A sophisticated world can still be a poor model of the intended use. | Require a fidelity statement; no claim of real-world impact from simulation alone.

 | Living labs | Users co-create innovations in real-world settings. [[10]](research.html#ref-10) | Ecological validity and repeatability pull in different directions. | Use synthetic conditions for comparison; later field validation requires separate protocols.

 | Assessment science / OSCE / ABET | Evidence-centred tasks, observed professional scenarios and engineering outcomes. [[12]](research.html#ref-12), [[13]](research.html#ref-13), [[14]](research.html#ref-14) | Rich observation can still support an invalid inference about competence. | Specify constructs, train raters and test transfer before hiring or certification use.

 | Chaos / fault injection | Hypotheses about steady state tested through disruptive events. [[15]](research.html#ref-15) | Breaking an implementation alone says little about selecting the right intervention. | Link stress families to declared assumptions, then run common held-out tests.

 | Reproducible research | Independent artifact evaluation and reproduction. [[16]](research.html#ref-16) | Availability of code is weaker than reproduced results. | Require independent run identity, manifest hashes and tolerance checks.

 | AI-agent benchmarks | Repository tasks, fresh instances and benchmark mutation. [[17]](research.html#ref-17), [[18]](research.html#ref-18), [[19]](research.html#ref-19) | Leakage, underspecified tests and benchmark-specific optimisation remain concerns. | Isolate scoring; use external observations, negative controls and fresh evaluation families.

 | AI-era hackathon redesign | HackerRank Orchestrate already centres AI-agent building and evaluation. [[28]](research.html#ref-28) | Allowing AI and adding evaluation is not differentiation. | Treat human–AI attribution as unresolved and test protocol value directly.


## Mechanism ledger


Names below identify concrete records and procedures. None is asserted to be patentable or unprecedented. The candidate contribution is their explicit linkage and the evidence needed to validate that linkage.

 | Mechanism | Prior art | Unresolved problem | Proposed mechanism | Test and novelty boundary

 | Decision receipt | Preregistration; evidence-centred assessment. [[12]](research.html#ref-12), [[21]](research.html#ref-21) | Final narratives can hide hindsight. | Timestamp a problem choice, alternatives, prediction, expected harms, budget and stopping rule before the intervention test. | Audit receipt order; compare blinded ratings with and without receipts. Not claimed novel in isolation.

 | Paired consequence record | Common random numbers and controlled comparisons. [[22]](research.html#ref-22) | A changing world makes before/after comparisons misleading. | Run frozen baseline and intervention against an identical exogenous schedule. Keep intervention-dependent responses causal in each branch. | Report seed-wise deltas, uncertainty and guardrail vectors. Coupling may fail; disclose divergence.

 | Revision trace | Hypothesis testing and engineering decision records. [[12]](research.html#ref-12), [[21]](research.html#ref-21) | A polished account cannot show when contrary evidence became available. | Link each amended prediction to signed evidence and the next action, including justified non-action. | Blind raters judge warranted revision, not number of pivots. Incremental validity unproven.

 | Assumption challenge pairing | Chaos engineering, held-out tasks and evolving environments. [[15]](research.html#ref-15), [[19]](research.html#ref-19), [[23]](research.html#ref-23) | Personalised stress can make participants incomparable. | Map declared dependencies to a public fault taxonomy; reserve a common holdout panel for ranked results. | Personalised faults are diagnostic only; measure transfer on the common panel.

 | Deferred-action test | Controlled experimentation; professional judgement assessment. [[12]](research.html#ref-12), [[14]](research.html#ref-14) | Build-only incentives penalise justified restraint. | Allow a no-change submission with evidence, predicted avoided harm and a falsifiable trigger for future action. | Baseline always runs. No automatic credit for abstention; review opportunity cost and rejected feasible choices.


## Research method


Research cut-off: 28 September 2026. Searches covered academic literature, official competition documentation, standards organisations, assessment bodies and technical documentation. Search families included hackathon outcomes, AI developer productivity, cyber-range assessment, adversarial scoring, scenario assessment, simulation validity, reproducibility, benchmark leakage, preregistration, environment co-evolution and standards governance.

Primary sources and research papers were preferred. Preprints are explicitly labelled. Company documentation establishes what a format claims to do; it does not independently establish validity. Search-result excerpts were sufficient for narrow landscape descriptions; linked documents support deeper inspection. This is not a systematic review, exhaustive patent search or independent replication.

Selection rule: retain sources that establish a close mechanism, challenge the central thesis, or change an implementation requirement. Reject promotional claims as evidence of superiority. An absence claim requires more than a missing search result. No empirical user data was collected for this draft.


## Adversarial review & open questions


 | Reviewer objection | Revision made | Still unresolved

 | CTO: “This is marketing for a simulator.” | Reduced the claim to a portable evidence contract; published all dependencies and the toy model’s limits. | Does the protocol change a real engineering decision?

 | Academic: “You have no construct validity.” | No certification, general ability score or human-only inference. Added blinded rating and transfer tests. | External validity, inter-rater reliability and domain bias.

 | Organiser: “This already exists.” | Named close precedents and dropped “first” and environment novelty claims. | Does the linked protocol add enough value to justify adoption?

 | Engineer: “I can game the metric.” | Added independent guardrails, matched baselines, negative controls, a frozen holdout and tamper boundaries. | Unknown proxies, simulator exploits and hidden distribution leakage.

 | AI researcher: “An agent can generate all your reasoning.” | Treat the team and its tools as the unit; receipts are commitments, not proof of cognition. | How to validate individual responsibility with increasingly autonomous agents?

 | Operator: “This is unaffordable.” | Capability profiles instead of prestige tiers; containers are sufficient for a narrow scenario. | Cost per valid replay, facilitator hours and repeated-run feasibility.

 | Participant: “The domain experts always win.” | Common briefing, practice environment, matched resources and separate domain reporting. | Accessibility, prior exposure and differential measurement effects.

Additional failure modes: teams deliberately fail to earn adaptation credit; explore too many outcomes then report only one; exfiltrate hidden seeds; alter telemetry; use external paid agents beyond declared budgets; discover the simulator’s numerical shortcuts; or optimise the mean at the expense of minority outcomes. Report attempted gaming and rejected evidence, and allow an independent appeal.


## References


Sources are not endorsements or affiliations. Dates and limitations below distinguish evidence from proposed methodology.

- [On Hackathons: A Multidisciplinary Literature Review ↗](https://doi.org/10.1145/3544548.3581234)CHI 2023 · peer-reviewed review; purposes, formats, processes and outcomes.
- [The Impact of AI on Developer Productivity: Evidence from GitHub Copilot ↗](https://arxiv.org/abs/2302.06590)Peng et al., 2023 · controlled task study; limited external generalisation.
- [Early-2025 AI and experienced open-source developer productivity ↗](https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf)METR, 2025 · randomised study in mature repositories.
- [We are Changing our Developer Productivity Experiment Design ↗](https://metr.org/blog/2026-02-24-uplift-update/)METR, February 2026 · selection effects limit newer estimates.
- [Cyber Grand Challenge ↗](https://www.darpa.mil/research/programs/cyber-grand-challenge)DARPA · autonomous discovery, repair and deployment precedent.
- [AI Cyber Challenge: final scoring guide announcement ↗](https://www.darpa.mil/news/2025/ai-cyber-challenge-scoring)DARPA, 2025 · functionality-preserving repair and automated scoring.
- [The Cyber Range: A Guide ↗](https://www.nist.gov/system/files/documents/2023/09/29/The%20Cyber%20Range_A%20Guide.pdf)NIST-hosted guide · simulated environments and workforce assessment.
- [Kaggle competition documentation ↗](https://www.kaggle.com/docs/competitions)Kaggle · measurable submissions and competition evaluation.
- [RoboCup Rescue ↗](https://robocup.de/en/major/rescue)RoboCup Germany · simulation and embodied rescue assessment.
- [Living Labs: origins, developments and future perspectives ↗](https://enoll.org/wp-content/uploads/2025/01/ENoLL-2025.-Living-Lab-origins-developments-and-future-perspectives.pdf)ENoLL, 2025 · co-creation in real-world innovation settings.
- [NASA Tournament Lab ↗](https://www.nasa.gov/directorates/stmd/prizes-challenges-crowdsourcing-program/center-of-excellence-for-collaborative-innovation-coeci/nasa-tournament-lab/)NASA · institutional challenge-based problem solving.
- [A Brief Introduction to Evidence-Centered Design ↗](https://www.ets.org/research/policy_research_reports/publications/report/2003/hsgs.html)Mislevy, Almond & Lukas, ETS, 2003 · claims, evidence and assessment tasks.
- [Engineering accreditation criteria 2025–2026 ↗](https://www.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2025-2026/)ABET · problem formulation, experimentation and engineering judgement.
- [A sample OSCE station ↗](https://www.gmc-uk.org/registration-and-licensing/join-our-registers/plab/plab-2-guide/a-sample-osce-station)General Medical Council · scenario-based professional assessment.
- [Principles of Chaos Engineering ↗](https://principlesofchaos.org/)Primary principles · steady-state hypotheses and realistic disruptive events.
- [Lessons from Five Years of Artifact Evaluation at EuroSys ↗](https://www.sigops.org/2025/lessons-from-five-years-of-artifact-evaluation-at-eurosys/)ACM SIGOPS, 2025 · independent artifact evaluation and reproduction.
- [SWE-bench ↗](https://www.swebench.com/)Official benchmark · repository-level software issue resolution.
- [SWE-rebench: decontaminated evaluation ↗](https://arxiv.org/abs/2505.20411)2025 preprint · fresh tasks and contamination concerns.
- [Saving SWE-Bench: A Benchmark Mutation Approach ↗](https://arxiv.org/abs/2510.08996)2025 preprint · benchmark mutation already exists.
- [Planning and Acting in Partially Observable Stochastic Domains ↗](https://www.cassandra.org/arc/papers/aij98.pdf)Kaelbling, Littman & Cassandra, 1998 · POMDP foundations.
- [Preregistration ↗](https://www.cos.io/initiatives/prereg)Center for Open Science · predictions distinguished from post-hoc explanation.
- [Bayesian Optimization Allowing for Common Random Numbers ↗](https://arxiv.org/abs/1910.09259)2019 preprint · shared stochastic conditions in comparison.
- [LLM-POET: Evolving Complex Environments ↗](https://arxiv.org/abs/2406.04663)2024 preprint · environment/agent co-evolution is prior art.
- [Digital Twins for Advanced Manufacturing ↗](https://www.nist.gov/programs-projects/digital-twins-advanced-manufacturing)NIST · validation, uncertainty and reference implementations.
- [Simulation interoperability standards products ↗](https://www.sisostandards.org/page/StandardsProducts)SISO · existing simulation interoperability standards.
- [Guide to the IETF standards process ↗](https://www.ietf.org/process/process/)IETF · open review and running implementations; no affiliation implied.
- [Conformity Assessment Basics ↗](https://www.nist.gov/standardsgov/conformity-assessment-basics)NIST · declaration, certification and accreditation are distinct.
- [Behind the Scenes of HackerRank Orchestrate ↗](https://www.hackerrank.com/blog/behind-the-scenes-of-hackerrank-orchestrate/)HackerRank, 2026 · AI-agent hackathon and evolving evaluation.
- [Multi-tenancy ↗](https://kubernetes.io/docs/concepts/security/multi-tenancy/)Kubernetes · namespace isolation requires additional controls.
- [OpenTelemetry signals ↗](https://opentelemetry.io/docs/concepts/signals/)OpenTelemetry · metrics, logs and traces.
- [Global Brand Database ↗](https://www.wipo.int/en/web/global-brand-database)WIPO · database coverage is incomplete; national searches also matter.
- [Patent subject matter eligibility ↗](https://www.uspto.gov/patents/laws/examination-policy/subject-matter-eligibility)USPTO · US eligibility guidance; not a patentability opinion.
- [Reducing overfitting in challenge-based competitions ↗](https://arxiv.org/abs/1607.00091)2016 preprint · adaptive leaderboard overfitting and Ladder precedent.

# Specification

## Status & scope


Research Draft 0.1 / proposed requirements

This is a proposed specification. MUST, SHOULD and MAY express intended conformance rules for this draft; they do not imply recognition by a standards body. No challenge or person is currently certified.

The assessment unit MUST be the team with its declared tools, time and resource envelope. The claim MUST be scoped to a scenario family and outcome. Scores MUST NOT be represented as general human ability or real-world deployment safety. Comparisons across different scenarios or profiles are prohibited without an equating study.

A host MUST publish the scenario manifest, construct map, scoring profile, guardrails, appeal procedure and permitted AI resources before the challenge. Hidden seeds and expected answers remain protected. Participants receive an environment and a mission boundary, rather than a prescribed solution. “Discover a problem” does not mean “guess the organiser’s secret problem.”


## Hackathon Lite


MERENIC is an evidence protocol for technical challenges; its Hackathon Lite profile applies it to time-boxed hackathons. This proposed entry profile works with an ordinary repository and a reproducible test. It does not require a simulated city or dedicated sandbox.

- Disclose at kickoff. Record the repository, pre-existing code, permitted tools, AI models or agents, and resource budget. Declare later tool changes as amendments.
- Commit a decision. Before the final test, timestamp the problem, rejected alternative, measurable prediction and failure condition. Attach the commit SHA and artifact hash.
- Submit a claim card. Link each claim to a test, retain failed and null results, disclose AI contributions, and name the conditions the result does not cover.
- Replay. A separate volunteer or CI operator follows the instructions and reports whether the claim reproduces within its declared tolerance.

Planning estimates: 10 minutes for disclosure, 15 for the receipt, 20 for the claim card and 10 for replay. These are unvalidated design estimates, not measured event overhead. Pilot organisers should record actual participant and judge time.

Lite produces a bounded, reproducible claim. It does not claim Foundation-profile conformance when paired environments, controlled disruptions or independent instrumentation are absent. Events should publish which requirements they implement, rather than presenting every profile as equivalent evidence.

### Conditions for a fair event

Publish the pre-existing-code rule, tool allowances, common budget, judge workload and accessible participation options before kickoff. Participants retain their existing intellectual property; any publication or reuse of submissions requires explicit terms and consent. Do not require disclosure of private prompts, personal data or credentials.

AI assistance may check schemas and organise evidence. Human reviewers assess relevance, tradeoffs, uncertainty and appeals. Disclose the model, prompt and role if automated review affects a decision; do not score the polish of generated prose as engineering evidence.


## Six foundational primitives


 | Primitive | Required meaning

 | Environment | A versioned world with a resettable initial state and a declared fidelity boundary.

 | Actors | Participants, users, services and agents; permissions and objectives are explicit.

 | Constraints | Resource budgets, safety boundaries, information access and non-negotiable limits.

 | Events | Scheduled or conditional changes with a reproducible trigger and bounded severity.

 | Telemetry | Externally collected observations with provenance, timestamps and known blind spots.

 | Outcomes | Selected effects, guardrails, uncertainty and independent reproduction conditions.


## Formal technical model


The state model adapts the existing partially observable decision-process formulation. It is not a new theory of state transitions. [[20]](research.html#ref-20)

Scenario C = (S, A, O, T, Z, μ₀, K, U, Ω, V)
s₀ ~ μ₀;  oₜ ~ Z(sₜ, visibilityₜ)
sₜ₊₁ = T(sₜ, aₜ, eₜ, ξₜ)
Rₜ = (observation_refs, hypothesis, prediction,
      alternatives, guardrails, budget, stopping_rule)
Eₜ = (manifest_hash, run_id, receipt_id, artifact_hash,
      event_schedule_hash, observations, outcomes, gaps)
Hₜ = (o₀, R₀, a₀, E₀, …, oₜ, Rₜ)
π(Hₜ) → aₜ or justified no-change

S is the full state; A includes observation requests, interventions and no-change; O is participant-visible information; T is the transition rule; Z defines visibility; μ₀ defines initial conditions. K contains constraints, U the declared outcome vector, Ω the held-out world families and V the versioned evaluator. Events eₜ and external randomness ξₜ are recorded separately from intervention-induced responses.

The proposed extension is an audit relation: every scored claim points to a pre-test receipt, admissible observations, an artifact and a comparison run. Amendments append records; they cannot rewrite history. An unexplained missing link makes the claim unverified, not necessarily false. A receipt stores a concise decision rationale, not private chain-of-thought.

### Information and time

Public: mission boundary, interfaces, resource costs, metric definitions, fault families and practice tasks. Hidden: future seeds, final fault instances and evaluator internals. Discoverable: bottlenecks, affected actors and causal dependencies through permitted observations. The host records wall-clock and simulation time, clock offsets and observation availability. Leaked information invalidates the affected comparison under the appeal policy.


## The evidence protocol

- Observe. Inspect a common initial world; record observation provenance and measurement gaps.
- Select. Submit up to three candidate problems and one selected claim. Explain relevance, uncertainty and opportunity cost. A domain expert checks admissibility without prescribing a solution.
- Commit. Timestamp a decision receipt before confirmatory testing. Exploratory runs stay labelled exploratory.
- Intervene. Submit an immutable artifact or explicit no-change action. Check constraints before activation.
- Compare. Run a frozen baseline and intervention on paired external schedules. Repeated runs use independent schedule seeds.
- Challenge. Run bounded assumption tests for diagnosis; score resilience only on the common frozen holdout.
- Revise. Link new evidence to amended predictions. Justified persistence and restraint are allowed.
- Reproduce. A separate operator resets the environment, executes the manifest and reports tolerance results.

Hosts MUST retain failed, cancelled and null-result runs. Hosts SHOULD publish a redacted evidence bundle and reviewer disagreements. Personally identifiable data and participant secrets MUST NOT appear in public artifacts.


## Measuring consequences


For a lower-is-better metric, on held-out world w and seed j:
  Δ(w,j) = Y_baseline(w,j) − Y_intervention(w,j)
  effect(w) = mean_j Δ(w,j)
Report: raw paired deltas, mean, interval, cost and guardrail vector.
Do not replace all of these with a single percentage.

The baseline MUST be independently controlled and frozen before interventions. The host MUST prevent participants from degrading it. Pair exogenous demand, not every random draw: intervention-induced events can legitimately differ. Use keyed random streams per actor/event so a changed code path does not silently shift every subsequent sample. If coupling assumptions fail, disclose that and use independently randomised replications.

Predeclare replications, practical effect threshold, aggregation and uncertainty method. A pilot may use a paired bootstrap interval across independent scenario seeds; time-correlated requests are not independent samples. For multiple selected claims, publish exploratory status or a multiplicity adjustment. Do not repeatedly inspect holdout feedback to choose a winner.

Guardrails are absolute eligibility conditions, not subtractive penalties that a good primary metric can outweigh. Examples: each district’s delay ceiling, zero prohibited network access and a hard compute budget. Exceeding one yields “constraint not met” for that run, with a reason and review path. Uncertainty crossing the minimum practical benefit yields “inconclusive,” not automatic success.


## Evaluation architecture


Machines verify schema, artifact identity, ordering, budgets, integrity and replay. The environment emits measured effects, distribution slices and failures; it does not decide their social value. Human experts assess relevance, reasoning, tradeoffs and limits through anchored rubrics. Do not score private cognition, personal worth, narrative polish, raw experiment count or unobserved real-world safety.

Two proposed profiles illustrate a construct-led weighting scheme. The discovery profile gives 50% to discovery, evidence and experiment design; the operational profile emphasises resilience and engineering. These are design hypotheses, not empirically optimised weights. Hosts choose and freeze one profile before participants begin.

 | Dimension | Discovery % | Operational % | Evidence anchor

 | Problem discovery | 20 | 8 | Relevance, competing explanations, selection justified by evidence

 | Evidence quality | 15 | 12 | Provenance, missingness and warranted claims

 | Systems reasoning | 10 | 10 | Dependencies, feedback and consequences beyond target

 | Experiment design | 15 | 8 | Controls, precommitment and uncertainty

 | Engineering quality | 5 | 12 | Correctness, operability and maintainability

 | Technical decisions | 10 | 8 | Tradeoffs, budget and rejected options

 | Resilience | 5 | 17 | Common held-out fault panel

 | Reproducibility | 10 | 10 | Independent replay within declared tolerance

 | Adaptation | 5 | 5 | Appropriate updates, including maintaining a justified decision

 | Measured outcome | 5 | 10 | Held-out improvement with guardrails satisfied

Human dimensions use a 0–4 scale: 0 absent or contradicted; 1 asserted with major gaps; 2 partially supported with relevant limitations; 3 supported with alternatives and appropriate limits; 4 independently corroborated and robust to a material counterexample. Dimension-specific examples are required before a scored edition.

At least two blinded raters score judgement dimensions independently before discussing disagreements. Report agreement and adjudication, not just consensus. Avoid counting the same evidence twice: discovery assesses selection, evidence assesses provenance, experiment design assesses identification. Machine dimensions need scenario-specific mappings from raw outcomes to 0–4 anchors, frozen in the manifest; there is no universal normalisation.

Eligibility first: constraints + integrity + required evidence.
If eligible: profile index = Σ weightᵢ × scoreᵢ / 4.
Always publish the dimension vector and uncertainty beside the index.
Missing required evidence → unverified, never imputed as success.

For Edition 01, retain profile indices for analysis; use broad evidence categories rather than precision rankings until rater reliability and holdout stability are examined. Sensitivity analysis must show whether plausible alternative weights reverse conclusions.


## Sandbox architecture


PARTICIPANT PLANE                   TRUSTED CONTROL PLANE
Workspace / permitted AI gateway →  Submission & receipt service
       ↓                             ↓ constraint admission
Per-team VM or isolated runtime ←  Scenario orchestrator
       ↕ APIs / synthetic actors     ↓ events / schedules
Instrumented world + intervention   Paired baseline runner
       ↓ external probes             ↓ independent probes
       └────────────→ Telemetry collector
                       ↓ append-only evidence store
                    Offline evaluator → human review → appeal
                       ↓
                    Independent replay worker

The control plane MUST be inaccessible to participant credentials. Collector ingestion credentials MUST NOT grant score-write access. External probes observe user-visible outcomes; application self-reported metrics are supporting evidence only. Store an append-only index with server signatures and object hashes in a separate account; a hash alone does not prove that an observation is true.

A Foundation implementation can use Docker Compose inside separate team VMs, an API simulator, a scheduler, an external collector and object storage. Kubernetes is optional. Namespaces alone do not constitute a security boundary for hostile workloads. [[29]](research.html#ref-29) OpenTelemetry supplies signal conventions; the protocol adds evidence identity and scenario metadata. [[30]](research.html#ref-30)

Advanced deployments can add Kubernetes jobs, snapshotting, traffic generators, fault controllers and synthetic agents. Frontier deployments may federate simulation engines or digital twins; a fidelity and uncertainty report remains mandatory. A real API needs a replay adapter or an explicit non-reproducible designation. Prefer recorded or simulated APIs for scored runs.

### Threat and operations boundary

Use per-team accounts, deny-by-default egress, quotas, non-privileged workloads, isolated storage and short-lived scoped credentials. Do not mount host control sockets. Use VMs or equivalent isolation for adversarial code. The operator needs kill switches, timeouts, reset tests, incident logs and clean teardown. A credential-compromise scenario uses synthetic credentials and in-sandbox resources only.

Fault records include family, trigger, duration, severity, affected dependency, expected invariant and recovery condition. Supported families include traffic increase, API delay, database loss, network partition, synthetic credential compromise, changed user demand, dependency failure, budget cuts, drift and simulated supply-chain incidents. Each must have a reason connected to an assumption; random destruction is not an assessment.

### Cost and scale

Estimated compute = teams × (interactive hours × workspace rate + world families × replications × paired-run hours × runner rate), plus storage, model usage and staff time. Rates are deployment-specific; no price is asserted here. Cache initial snapshots, queue short jobs and share immutable scenario assets. Cap API spend, log model/tool versions and report resource differences. No cross-team comparisons across unequal budgets without explicit stratification.


## Implementation profiles


 | Profile | Required capability | What it does not imply

 | Foundation | Resettable environment, timestamped receipts, externally measured baseline/intervention pair, bounded event, exportable evidence, independent replay. | No Kubernetes or digital twin requirement. Manual orchestration may conform.

 | Advanced | Foundation plus automated paired runs, event schedules, trace correlation and isolated evaluation services. | Automation does not confer greater assessment validity.

 | Frontier | Advanced plus validated domain simulation, multi-agent dynamics or federated worlds and uncertainty reporting. | Fidelity is a documented claim, not a prestige badge.

Profiles describe implementation capabilities, not human ability levels. The tiering is a proposal informed by the diversity of existing simulation and assessment infrastructure, not an established maturity model. Conformance requires the same evidence contract at every level. An institution may implement the contract without buying CyberMindSpace technology.


## Adoption artifacts


The downloadable manifest and receipt schema are starting points for implementers. They are not a production conformance suite. Before Edition 01, a host must add a real scenario adapter, integrity service, dimension anchors, reset tests and an independent replay harness.

[Example scenario manifest ↓](assets/scenario.example.json)[Decision receipt JSON Schema ↓](assets/receipt.schema.json)[Download this specification as HTML ↓](assets/specification.html)[Research and adversarial review dossier ↓](assets/research-dossier.md)

# Governance and validation

## Roles and relationships


MERENIC is a research initiative initiated by Majji Pradeep Kumar and developed with support from CyberMindSpace Private Limited. At Research Draft 0.1 there is no independent standards body, charter or appointed committee. The table separates the roles that exist now from those that are proposed.

 | Role | Who | Responsibility | Limit

 | MERENIC methodology and research | The published thesis, protocol and study plan | Defines claims, evidence requirements and evaluation methods; revised through public review and field evidence. | Not a recognised standard or a validated assessment instrument.

 | Initiator | Majji Pradeep Kumar | Maintains the research draft, coordinates review and prepares Edition 01. | Proposes changes through the decision process below, with rationale published.

 | Supporting organisation | [CyberMindSpace Private Limited](https://cybermindspace.com/) | Supports development and the reference implementation. | Does not decide research conclusions on its own. MERENIC is not currently institutionally independent of CyberMindSpace.

 | Research Advisory Committee | Not yet formed | Would review methodology, drafts and field results, and publish its dispositions. | Members are listed only after they explicitly agree to participate.

 | Edition organisers | Named for each edition | Run an edition under a pinned specification, scenario and evaluator version. | Cannot change scoring rules during an edition.

 | Sponsors and partners | None announced | May provide funding, infrastructure or venues. | No authority over research findings, participant outcomes or assessment results.

Principle

Financial or infrastructure support does not confer authority over assessment outcomes, research conclusions or protocol revisions.

No chair or steward role is assigned: such roles require an adopted charter and appointment process. The governance below is proposed, not a claim that an independent standards body or committee already exists. Institutional credibility must come from published decisions and review, not titles.


## Research Advisory Committee


The proposed Research Advisory Committee would combine senior engineering and technology leadership with assessment research, domain operations and independent implementer perspectives. Seats remain unconfirmed. C-suite seniority alone is not a substitute for methodological expertise.

No members have been appointed. Members will be listed on the About page only after they explicitly agree to participate; affiliations will be shown for identification and will not imply institutional endorsement.

Members would review the thesis, construct map, architecture, failure taxonomy, successive drafts and field results. Each review produces a public disposition: accept, revise or reject, with rationale and conflicts disclosed. Members should challenge commercial assumptions as well as technical ones, including those of CyberMindSpace.

Proposed terms: one-year renewable appointments, a public interests register, recusal from conflicts, and an independent methodology lead. The first charter should set quorum at two-thirds of filled seats and require a two-thirds majority of non-conflicted voting members for a release, with minority opinions recorded. These rules take effect only after charter ratification.


## Decision process


An editor maintains the draft; a methodology group reviews assessment claims; implementers report interoperability defects; the committee reviews disputed changes and release readiness. CyberMindSpace supports development and the reference implementation but should not be the sole judge of its own conformance.

Each proposal carries a problem statement, affected requirements, evidence, compatibility impact and review window. Proposed public comment window: 30 days for substantive drafts. Urgent integrity fixes may be released sooner with written rationale and retrospective review. An appeal panel must exclude original scorers and conflicts.

This takes inspiration from open review and implementation-led standards development. It does not copy the IETF’s status labels or imply IETF sponsorship. [[26]](research.html#ref-26)


## Version lifecycle


 | Stage | Status | Exit evidence

 | 0.1 · Research Draft | Current · 28 Sep 2026 | Prior-art map, proposed contract, limitations and candidate experiment plan.

 | 0.2 · Review Draft | Planned | Charter adopted; advisors confirmed; comment dispositions and refined rubric.

 | 0.5 · Implementer Draft | Planned | Two independently operated adapters exchange evidence; reset and replay checks pass.

 | Edition 01 · Pilot | In development | Frozen protocol and registered analysis plan; study and operations approvals completed where applicable.

 | 0.x · Evidence Revision | Planned | Publish null results, disagreements, failure modes, costs and changes.

 | 1.0 · Stable Specification | Not scheduled | Independent implementations, acceptable validation evidence and governance release decision.

Edition numbers identify field implementations; specification versions identify compatible requirements. An edition pins a specification, scenario and evaluator version. Never change scoring rules mid-edition without marking the affected runs incomparable and offering reruns. Breaking changes increment the major version after 1.0; historical results retain their original evaluator.


## Edition 01 as an experiment


FIRST CONTROLLED PILOT · IN DEVELOPMENT

Edition 01 will be the first implementation of MERENIC. It will test the framework as well as its participants: alongside participants’ engineering judgement, it is designed to generate evidence about the protocol itself. The following is a proposed study plan; it is not a claim of preregistration or a completed pilot. No date, venue, organisers or sponsors have been announced, and the event will be documented separately from this research site.

 | Hypothesis | Primary observable | Falsification / revision trigger

 | H1: receipts improve judgement consistency. | Two blinded raters score full bundles and independently assigned artifact-only packets. | No improvement in agreement, or differences explained by extra reading time alone.

 | H2: paired evidence predicts held-out consequences better than presentation ratings. | Compare association with a separate common holdout result. | No improvement, wide inconclusive intervals or weight-sensitive conclusions.

 | H3: justified restraint can be assessed without rewarding inactivity. | Blinded review of selected no-change and intervention claims; matched scenario opportunity. | Raters reward non-action independent of evidence or cannot distinguish defensible cases.

 | H4: the protocol is portable. | Independent operators replay the same bundles against pinned manifests. | Outcome differences exceed declared tolerance or missing artifacts prevent replay.

 | H5: evidence overhead is operationally tolerable. | Time spent documenting, cost per valid run, completion and accessibility feedback. | Overhead exceeds the host’s preregistered limit or excludes participant groups.

### Proposed pilot design

Recruit a feasibility cohort; a provisional planning range is 12–20 teams, not a powered confirmatory sample. Stratify by relevant experience and provide the same orientation and resource limits. Randomise review packets across raters and counterbalance scenario order. Keep final holdout outcomes unavailable to reviewers. Analyse team-level units; do not pretend hundreds of telemetry points are hundreds of participants.

Before recruitment, freeze hypotheses, exclusion rules, primary endpoints, confidence intervals, rater agreement method, missing-data handling, budget ceiling and go/no-go thresholds. An assessment researcher should estimate a confirmatory sample using pilot variance. Report uncertainty and effect sizes; do not infer validity from a small pilot’s p-value.

Capture consented decision receipts, interventions, event schedules, tool/resource declarations, rubric judgements, errors and operational effort. Avoid blanket capture of private prompts or secrets. Publish de-identified artifacts only with explicit rights and retention terms. An independent reviewer examines adverse findings before the next release.


## Open adoption. Earned defensibility.


The recommended separation is an openly adoptable protocol and public conformance tests, alongside CyberMindSpace-operated scenario infrastructure, orchestration, support and validated environment libraries. Licensing is a governance decision still to be adopted; this draft does not silently assign ownership or claim an open-source licence on behalf of the organisation.

Potential defensibility comes from reliable adapters, scenario quality, independent field evidence, institutional trust and consented datasets—not the abstraction of observing and iterating. Hidden test sets need rotation and access controls, while score definitions and appeal rules stay public. Data advantage must be earned through permission, quality and representativeness.

Certification, if pursued, needs a separate impartiality and competency model. A host’s declaration is not third-party certification, and certification is not accreditation. [[27]](research.html#ref-27)

The technical implementation of deterministic branching, signed provenance or isolation might merit professional prior-art review once concrete code exists. This review establishes no patentability. Patent eligibility is jurisdiction-specific; US guidance treats abstract-idea questions separately from merely implementing a process in software. [[32]](research.html#ref-32)


## Changelog


1 October 2026 · Site update. Clarified the initiator and the supporting role of CyberMindSpace Private Limited; replaced “founding steward” wording with a roles table; added How It Works and About pages. Protocol requirements are unchanged, and Research Draft 0.1 remains dated September 2026.

30 September 2026 · Draft refinement. Added the proposed Hackathon Lite profile, tool and commit provenance fields, and an illustrated city/infrastructure introduction. Removed unconfirmed advisor placeholders. No field-validation claim is added.

0.1 · 28 September 2026. Initial research draft. Narrowed the AI productivity claim; acknowledged cyber-range, competition and assessment precedents; adopted a POMDP-based model; separated diagnostic stress from comparable holdout evaluation; removed human-only attribution claims; added evidence gates, no-change submissions and the Edition 01 validation plan.

No prior released versions, advisor approvals, completed field studies or certified implementations are claimed.


# Naming decision record

Twenty-four candidates were assessed qualitatively against conceptual depth, memorability, seriousness, pronunciation, visual identity, extensibility, thesis fit and collision risk. Common-word eliminations are linguistic judgement, not verified company searches. Shortlisted and high-fit candidates received indexed web checks.

- **MORPH**: Strong change metaphor, easy and visual; too broad and product-like; screen out.
- **MESH**: Systems relevance and pronunciation strong; descriptive and crowded; screen out.
- **MATRIX**: Extensible and serious but generic technical term; weak distinctiveness.
- **METHOD**: Institutional and clear; descriptive rather than ownable identity.
- **MOLT**: Memorable transition metaphor; biological and startup-like tone.
- **METRON**: Measurement fit excellent; actual simulation/evaluation company conflict.
- **MERIDIAN**: Institutional, global and visual; actual Meridian AI Standard conflict.
- **METHEXIS**: Participation concept and formal tone; pronunciation friction; existing spatial-AI product. Alternative with high conflict risk.
- **METRION**: Measurement association; existing platform and technical uses.
- **METRAVA**: Easy to say and extend; actual infrastructure and app conflicts.
- **METRIVEN**: Good evidence association; actual metrics-driven competition platform conflict.
- **MODULUS**: Strong technical meaning; descriptive mathematical term, less distinctive.
- **MIMESIS**: Simulation association; does not foreground intervention or evidence.
- **MORPHIC**: Change-oriented; sounds like a technology product rather than a specification.
- **MANIFOLD**: Systems depth and visual potential; generic technical usage and pronunciation ambiguity.
- **MOTIVE**: Problem intent; insufficient relation to evidence, broad word.
- **MARGIN**: Robustness margin; easily understood but too generic.
- **MANTLE**: Institutional feel; weak assessment relation, metaphor needs explanation.
- **METRIC**: Explicit measurement; overly descriptive and encourages score reduction.
- **MEASURE**: Direct thesis link; generic verb, awkward edition language.
- **MERALITH**: Coined, stable, serious, extensible; weaker evidence link; alternative pending clearance.
- **METRACLINE**: Technical but difficult pronunciation; indexed cleaning-product collision.
- **MERENIC**: Coined, compact, pronounceable with guidance; strong wordmark and institutional extensibility. Selected; no guaranteed clearance.
- **METHRON**: Measurement/method suggestion but close to Metron and harder to hear correctly.

Primary recommendation: MERENIC. Alternatives: MERALITH and METHEXIS (the latter has a substantive known technology conflict). A recommendation is not clearance. Exact-name MERENIC searches and indexed GitHub/LinkedIn/trademark searches did not reveal an obvious technical initiative in this review, but personal-name occurrences exist. Domain and handle availability are unverified; none was acquired. Search strings included name + technology/company/standard/trademark/domain and exact quoted names. WIPO guidance was reviewed; no complete national registry or similar-mark search was completed.

## Logo exploration

Three geometric directions were considered: a closed boundary with one displaced segment (boundary/intervention), concentric broken loops (feedback), and paired state bars (comparison). Broken loops became ambiguous at favicon scale; paired bars looked like a pause control. The revised identity uses a solid geometric letter M with a deep central fold. Final SVG geometry is the master; the wordmark SVG uses an Arial text fallback rather than licensed outlined type.

## Visual provenance

Two original photorealistic 3D illustrations were generated for this project: a city with a closed bridge and alternate route, and enterprise infrastructure with database failure and bounded-queue failover. It is conceptual artwork, not a photograph of an existing system. Canvas network animation is functional model visualisation. All output metrics in the interactive example arise from the bundled transparent toy model; none is an Edition 01 result.

## Delivery boundary

Implemented: responsive multi-page website, original image and logo files, research/thesis, formal draft, evaluation profiles, governance and study plan, JSON adoption artifacts, deterministic interactive example and export. Designed but not built: isolated production cyber range, trusted receipt/signing infrastructure, secure automated scorer and independent replay service. Pending external work: name clearance, advisor appointments, licensing adoption, preregistration, validation and pilot execution.
