Status & scope
This is a proposed specification. MUST, SHOULD and MAY express intended conformance rules for this draft; they do not imply recognition by a standards body. No challenge or person is currently certified.
The assessment unit MUST be the team with its declared tools, time and resource envelope. The claim MUST be scoped to a scenario family and outcome. Scores MUST NOT be represented as general human ability or real-world deployment safety. Comparisons across different scenarios or profiles are prohibited without an equating study.
A host MUST publish the scenario manifest, construct map, scoring profile, guardrails, appeal procedure and permitted AI resources before the challenge. Hidden seeds and expected answers remain protected. Participants receive an environment and a mission boundary, rather than a prescribed solution. “Discover a problem” does not mean “guess the organiser’s secret problem.”
Hackathon Lite
MERENIC is an evidence protocol for technical challenges; its Hackathon Lite profile applies it to time-boxed hackathons. This proposed entry profile works with an ordinary repository and a reproducible test. It does not require a simulated city or dedicated sandbox.
- Disclose at kickoff. Record the repository, pre-existing code, permitted tools, AI models or agents, and resource budget. Declare later tool changes as amendments.
- Commit a decision. Before the final test, timestamp the problem, rejected alternative, measurable prediction and failure condition. Attach the commit SHA and artifact hash.
- Submit a claim card. Link each claim to a test, retain failed and null results, disclose AI contributions, and name the conditions the result does not cover.
- Replay. A separate volunteer or CI operator follows the instructions and reports whether the claim reproduces within its declared tolerance.
Planning estimates: 10 minutes for disclosure, 15 for the receipt, 20 for the claim card and 10 for replay. These are unvalidated design estimates, not measured event overhead. Pilot organisers should record actual participant and judge time.
Lite produces a bounded, reproducible claim. It does not claim Foundation-profile conformance when paired environments, controlled disruptions or independent instrumentation are absent. Events should publish which requirements they implement, rather than presenting every profile as equivalent evidence.
Conditions for a fair event
Publish the pre-existing-code rule, tool allowances, common budget, judge workload and accessible participation options before kickoff. Participants retain their existing intellectual property; any publication or reuse of submissions requires explicit terms and consent. Do not require disclosure of private prompts, personal data or credentials.
AI assistance may check schemas and organise evidence. Human reviewers assess relevance, tradeoffs, uncertainty and appeals. Disclose the model, prompt and role if automated review affects a decision; do not score the polish of generated prose as engineering evidence.
Six foundational primitives
| Primitive | Required meaning |
|---|---|
| Environment | A versioned world with a resettable initial state and a declared fidelity boundary. |
| Actors | Participants, users, services and agents; permissions and objectives are explicit. |
| Constraints | Resource budgets, safety boundaries, information access and non-negotiable limits. |
| Events | Scheduled or conditional changes with a reproducible trigger and bounded severity. |
| Telemetry | Externally collected observations with provenance, timestamps and known blind spots. |
| Outcomes | Selected effects, guardrails, uncertainty and independent reproduction conditions. |
Formal technical model
The state model adapts the existing partially observable decision-process formulation. It is not a new theory of state transitions. [20]
Scenario C = (S, A, O, T, Z, μ₀, K, U, Ω, V)
s₀ ~ μ₀; oₜ ~ Z(sₜ, visibilityₜ)
sₜ₊₁ = T(sₜ, aₜ, eₜ, ξₜ)
Rₜ = (observation_refs, hypothesis, prediction,
alternatives, guardrails, budget, stopping_rule)
Eₜ = (manifest_hash, run_id, receipt_id, artifact_hash,
event_schedule_hash, observations, outcomes, gaps)
Hₜ = (o₀, R₀, a₀, E₀, …, oₜ, Rₜ)
π(Hₜ) → aₜ or justified no-changeS is the full state; A includes observation requests, interventions and no-change; O is participant-visible information; T is the transition rule; Z defines visibility; μ₀ defines initial conditions. K contains constraints, U the declared outcome vector, Ω the held-out world families and V the versioned evaluator. Events eₜ and external randomness ξₜ are recorded separately from intervention-induced responses.
The proposed extension is an audit relation: every scored claim points to a pre-test receipt, admissible observations, an artifact and a comparison run. Amendments append records; they cannot rewrite history. An unexplained missing link makes the claim unverified, not necessarily false. A receipt stores a concise decision rationale, not private chain-of-thought.
Information and time
Public: mission boundary, interfaces, resource costs, metric definitions, fault families and practice tasks. Hidden: future seeds, final fault instances and evaluator internals. Discoverable: bottlenecks, affected actors and causal dependencies through permitted observations. The host records wall-clock and simulation time, clock offsets and observation availability. Leaked information invalidates the affected comparison under the appeal policy.
The evidence protocol
- Observe. Inspect a common initial world; record observation provenance and measurement gaps.
- Select. Submit up to three candidate problems and one selected claim. Explain relevance, uncertainty and opportunity cost. A domain expert checks admissibility without prescribing a solution.
- Commit. Timestamp a decision receipt before confirmatory testing. Exploratory runs stay labelled exploratory.
- Intervene. Submit an immutable artifact or explicit no-change action. Check constraints before activation.
- Compare. Run a frozen baseline and intervention on paired external schedules. Repeated runs use independent schedule seeds.
- Challenge. Run bounded assumption tests for diagnosis; score resilience only on the common frozen holdout.
- Revise. Link new evidence to amended predictions. Justified persistence and restraint are allowed.
- Reproduce. A separate operator resets the environment, executes the manifest and reports tolerance results.
Hosts MUST retain failed, cancelled and null-result runs. Hosts SHOULD publish a redacted evidence bundle and reviewer disagreements. Personally identifiable data and participant secrets MUST NOT appear in public artifacts.
Measuring consequences
For a lower-is-better metric, on held-out world w and seed j: Δ(w,j) = Y_baseline(w,j) − Y_intervention(w,j) effect(w) = mean_j Δ(w,j) Report: raw paired deltas, mean, interval, cost and guardrail vector. Do not replace all of these with a single percentage.
The baseline MUST be independently controlled and frozen before interventions. The host MUST prevent participants from degrading it. Pair exogenous demand, not every random draw: intervention-induced events can legitimately differ. Use keyed random streams per actor/event so a changed code path does not silently shift every subsequent sample. If coupling assumptions fail, disclose that and use independently randomised replications.
Predeclare replications, practical effect threshold, aggregation and uncertainty method. A pilot may use a paired bootstrap interval across independent scenario seeds; time-correlated requests are not independent samples. For multiple selected claims, publish exploratory status or a multiplicity adjustment. Do not repeatedly inspect holdout feedback to choose a winner.
Guardrails are absolute eligibility conditions, not subtractive penalties that a good primary metric can outweigh. Examples: each district’s delay ceiling, zero prohibited network access and a hard compute budget. Exceeding one yields “constraint not met” for that run, with a reason and review path. Uncertainty crossing the minimum practical benefit yields “inconclusive,” not automatic success.
Evaluation architecture
Machines verify schema, artifact identity, ordering, budgets, integrity and replay. The environment emits measured effects, distribution slices and failures; it does not decide their social value. Human experts assess relevance, reasoning, tradeoffs and limits through anchored rubrics. Do not score private cognition, personal worth, narrative polish, raw experiment count or unobserved real-world safety.
Two proposed profiles illustrate a construct-led weighting scheme. The discovery profile gives 50% to discovery, evidence and experiment design; the operational profile emphasises resilience and engineering. These are design hypotheses, not empirically optimised weights. Hosts choose and freeze one profile before participants begin.
| Dimension | Discovery % | Operational % | Evidence anchor |
|---|---|---|---|
| Problem discovery | 20 | 8 | Relevance, competing explanations, selection justified by evidence |
| Evidence quality | 15 | 12 | Provenance, missingness and warranted claims |
| Systems reasoning | 10 | 10 | Dependencies, feedback and consequences beyond target |
| Experiment design | 15 | 8 | Controls, precommitment and uncertainty |
| Engineering quality | 5 | 12 | Correctness, operability and maintainability |
| Technical decisions | 10 | 8 | Tradeoffs, budget and rejected options |
| Resilience | 5 | 17 | Common held-out fault panel |
| Reproducibility | 10 | 10 | Independent replay within declared tolerance |
| Adaptation | 5 | 5 | Appropriate updates, including maintaining a justified decision |
| Measured outcome | 5 | 10 | Held-out improvement with guardrails satisfied |
Human dimensions use a 0–4 scale: 0 absent or contradicted; 1 asserted with major gaps; 2 partially supported with relevant limitations; 3 supported with alternatives and appropriate limits; 4 independently corroborated and robust to a material counterexample. Dimension-specific examples are required before a scored edition.
At least two blinded raters score judgement dimensions independently before discussing disagreements. Report agreement and adjudication, not just consensus. Avoid counting the same evidence twice: discovery assesses selection, evidence assesses provenance, experiment design assesses identification. Machine dimensions need scenario-specific mappings from raw outcomes to 0–4 anchors, frozen in the manifest; there is no universal normalisation.
Eligibility first: constraints + integrity + required evidence. If eligible: profile index = Σ weightᵢ × scoreᵢ / 4. Always publish the dimension vector and uncertainty beside the index. Missing required evidence → unverified, never imputed as success.
For Edition 01, retain profile indices for analysis; use broad evidence categories rather than precision rankings until rater reliability and holdout stability are examined. Sensitivity analysis must show whether plausible alternative weights reverse conclusions.
Sandbox architecture
PARTICIPANT PLANE TRUSTED CONTROL PLANE
Workspace / permitted AI gateway → Submission & receipt service
↓ ↓ constraint admission
Per-team VM or isolated runtime ← Scenario orchestrator
↕ APIs / synthetic actors ↓ events / schedules
Instrumented world + intervention Paired baseline runner
↓ external probes ↓ independent probes
└────────────→ Telemetry collector
↓ append-only evidence store
Offline evaluator → human review → appeal
↓
Independent replay workerThe control plane MUST be inaccessible to participant credentials. Collector ingestion credentials MUST NOT grant score-write access. External probes observe user-visible outcomes; application self-reported metrics are supporting evidence only. Store an append-only index with server signatures and object hashes in a separate account; a hash alone does not prove that an observation is true.
A Foundation implementation can use Docker Compose inside separate team VMs, an API simulator, a scheduler, an external collector and object storage. Kubernetes is optional. Namespaces alone do not constitute a security boundary for hostile workloads. [29] OpenTelemetry supplies signal conventions; the protocol adds evidence identity and scenario metadata. [30]
Advanced deployments can add Kubernetes jobs, snapshotting, traffic generators, fault controllers and synthetic agents. Frontier deployments may federate simulation engines or digital twins; a fidelity and uncertainty report remains mandatory. A real API needs a replay adapter or an explicit non-reproducible designation. Prefer recorded or simulated APIs for scored runs.
Threat and operations boundary
Use per-team accounts, deny-by-default egress, quotas, non-privileged workloads, isolated storage and short-lived scoped credentials. Do not mount host control sockets. Use VMs or equivalent isolation for adversarial code. The operator needs kill switches, timeouts, reset tests, incident logs and clean teardown. A credential-compromise scenario uses synthetic credentials and in-sandbox resources only.
Fault records include family, trigger, duration, severity, affected dependency, expected invariant and recovery condition. Supported families include traffic increase, API delay, database loss, network partition, synthetic credential compromise, changed user demand, dependency failure, budget cuts, drift and simulated supply-chain incidents. Each must have a reason connected to an assumption; random destruction is not an assessment.
Cost and scale
Estimated compute = teams × (interactive hours × workspace rate + world families × replications × paired-run hours × runner rate), plus storage, model usage and staff time. Rates are deployment-specific; no price is asserted here. Cache initial snapshots, queue short jobs and share immutable scenario assets. Cap API spend, log model/tool versions and report resource differences. No cross-team comparisons across unequal budgets without explicit stratification.
Implementation profiles
| Profile | Required capability | What it does not imply |
|---|---|---|
| Foundation | Resettable environment, timestamped receipts, externally measured baseline/intervention pair, bounded event, exportable evidence, independent replay. | No Kubernetes or digital twin requirement. Manual orchestration may conform. |
| Advanced | Foundation plus automated paired runs, event schedules, trace correlation and isolated evaluation services. | Automation does not confer greater assessment validity. |
| Frontier | Advanced plus validated domain simulation, multi-agent dynamics or federated worlds and uncertainty reporting. | Fidelity is a documented claim, not a prestige badge. |
Profiles describe implementation capabilities, not human ability levels. The tiering is a proposal informed by the diversity of existing simulation and assessment infrastructure, not an established maturity model. Conformance requires the same evidence contract at every level. An institution may implement the contract without buying CyberMindSpace technology.
Adoption artifacts
The downloadable manifest and receipt schema are starting points for implementers. They are not a production conformance suite. Before Edition 01, a host must add a real scenario adapter, integrity service, dimension anchors, reset tests and an independent replay harness.