From a running system
to evidence that holds.
What a team actually does in a MERENIC challenge, in ten steps. This page describes the participant experience; the formal requirements are in the Protocol.
Problem statement → build → demo. The final artifact carries most of the result.
Running environment → decisions → measured consequences → replay. The judgement behind the artifact becomes evidence.
Ten steps, four phases.
- 01Enter
Receive the environment
Teams receive a running system, the observations available to them, the actions they are permitted to take, their constraints and resources, and the assessment rules.
They are not handed a conventional problem statement and asked to build a solution. Working out what the problem is forms part of the task.
- 02Observe
Investigate the system
Teams inspect behaviour, telemetry, users, constraints, failure signals and any other evidence the environment exposes. What they looked at, and what they could not see, is recorded.
- 03Select
Choose what matters
The team identifies what it believes is a consequential problem or opportunity for intervention. Problem selection itself becomes part of the evidence.
- 04Commit
Record the decision before acting
Before intervening, the team records:
- what it believes is happening
- what it intends to change
- the consequence it expects
- its assumptions
- the trade-offs it is accepting
- the evidence that would prove it wrong
- 05Intervene
Change the system
Teams implement their intervention. AI tools are permitted: the protocol does not try to measure engineering ability by removing the tools engineers actually use. Teams declare the tools they used.
- 06Compare
Measure consequences
The relevant system state is compared before and after the intervention, under matched conditions. The question is not only “did the implementation work?” It is “what changed because of the engineering decision?”
Paired consequence record
- 07Challenge
Introduce a withheld condition
A condition the team did not optimise for is introduced: a traffic spike, a dependency failure, adversarial input, a resource constraint or changed user behaviour. The purpose is to expose assumptions and test whether the intervention remains defensible.
- 08Revise
Respond to evidence
Teams may revise, replace or defend their original intervention. How their reasoning changes, or why it holds, becomes evidence. A change of direction earns no automatic credit.
Revision trace - 09Reproduce
Replay the result
Where the scenario permits, the intervention and its claimed consequences are replayed independently, by someone other than the team, within declared tolerances.
- 10Assess
Evaluate the engineering judgement
Reviewers examine the evidence as a whole: problem selection, reasoning, prediction quality, the intervention, measured consequences, trade-offs, the response to the challenge, the quality of revision and reproducibility. The final artifact alone does not determine the result.
A queue under pressure.
An infrastructure and reliability scenario, shown step by step. It is illustrative, not a result from a MERENIC edition.

A service experiences increasing request latency and a growing queue.
The team finds that workers are saturated and decides the queue is the most consequential bottleneck.
Before acting, the team predicts that increasing worker concurrency will reduce queue depth without materially increasing database pressure.
Worker concurrency is increased.
Queue depth improves, but database connection saturation rises significantly. Part of the prediction is already contradicted.
A traffic spike is introduced. The database becomes the dominant bottleneck, and retries amplify the failure.
The team revises its approach, for example with bounded concurrency, backpressure and retry limits, or another defensible architecture, and records why.
Reviewers do not simply ask whether the final system works. They examine why the original problem was selected, whether the prediction was reasonable, what evidence contradicted it, whether the team recognised the second-order effect, how well it revised, and whether the final result can be reproduced.
MERENIC is not a recognised standard or a validated assessment instrument. Edition 01, the first controlled pilot, will test the framework as well as its participants.