Scoring code you have not run is guesswork. So TheQuizMaster runs it — real candidate submissions, fourteen languages, on infrastructure I own and pay for.
Which means accepting hostile code as a normal input and designing for it, not as an edge case.
The threat model is not what people expect
Ask an engineer what worries them about running untrusted code and they will say container escape. It is the interesting answer, and for most of us it is the wrong one to spend the budget on.
Container escapes need a kernel vulnerability, are patched quickly, and are overwhelming rarer than the thing that will actually happen to you: someone submits an accidental infinite loop, or a recursion that eats the heap, or a program that writes to disk in a tight loop. There is no malice involved and the effect on every other tenant is the same as an attack.
Ordered by how often they actually occur:
- Resource exhaustion. Infinite loops, fork bombs, memory allocation in a loop, filling the disk. Most are accidents.
- Data exfiltration through the network. A submission that posts the test fixtures to a remote endpoint, or fetches an answer from one.
- Persistence between runs. Writing to a shared volume that the next submission can read, which turns an assessment platform into a message bus between candidates.
- Escape. Real, worth defending against in depth, and last on this list because the first three will happen this week.
Designing for the top of the list gets you most of the safety for most of the cost.
The limits that do the work
Every one of these is boring and every one has earned its place.
| Limit | Why |
|---|---|
| Wall-clock timeout | The only defence against an infinite loop. Kills the process group, not just the process. |
| CPU time limit | Separate from wall clock — catches a busy loop that would otherwise spend its whole budget spinning. |
| Memory cap | An OOM kill of one container is fine. An OOM kill chosen by the host is not. |
| Process limit | Without a PID cap, a fork bomb takes the host down regardless of every other limit. |
| Open file descriptors | Cheap to set, prevents a whole class of resource starvation. |
| No network | Removes exfiltration and "download the answer" in one line. |
| Read-only root filesystem | Plus a small tmpfs for scratch. Bounded and gone at exit. |
| Non-root user | Table stakes, and still frequently missed. |
| Dropped capabilities | Drop all, add none. Compilers do not need CAP_NET_ADMIN. |
Two of these deserve emphasis because they are the ones most often set wrong.
The timeout must kill the process group. A submission that spawns a child and exits leaves the child running. If your supervisor waits on the parent PID only, the container keeps consuming CPU after you believe the run is finished, and the leak accumulates across submissions until the host degrades.
--pids-limit is not optional. Memory limits do not stop a fork bomb —
each process is small. The PID cap is the only thing that does, and it is one
flag.
No network means no network
This is the highest-value single decision, and the temptation to soften it is constant. Someone will want to allow a package installer. Someone will want to let a submission call a documentation API.
The moment there is an allowlist, there is an allowlist bug. Fully offline is a property you can state and verify; partially offline is a configuration you have to audit forever.
Everything a submission needs is baked into the image ahead of time: the runtime, the standard library, and whichever packages the exercise legitimately requires. The image is built once, pinned, and scanned. Dependency resolution happens at build time, under my control, never at run time under the candidate's.
The practical consequence is a family of images rather than one, and a build pipeline that keeps them current. That is a real cost. It is smaller than the cost of reasoning about egress rules across fourteen language ecosystems.
Fourteen languages, one contract
The thing that keeps this manageable is refusing to special-case per language. Every runner implements the same contract:
{
"language": "java",
"sourceFiles": [{ "path": "Solution.java", "content": "..." }],
"tests": [{ "path": "SolutionTest.java", "content": "..." }],
"limits": { "wallClockMs": 10000, "memoryMb": 512, "pids": 64 }
}{
"outcome": "COMPLETED",
"compile": { "ok": true, "output": "" },
"tests": { "passed": 12, "failed": 0, "durationMs": 840 },
"truncatedOutput": false
}The differences between languages — compile step or not, how a test runner reports, how the toolchain writes to stderr — are absorbed inside each image behind that contract. The orchestrator does not know what Java is. It knows how to start a container, enforce limits, and read a result document.
That is what makes adding a fifteenth language a day of work rather than a refactor.
Output is an attack surface too
A submission that prints in a loop will produce gigabytes of stdout, and if your supervisor is accumulating that in memory before writing it anywhere, the submission has just taken down the orchestrator rather than its own container.
Cap the captured output in bytes, truncate with a marker, and treat exceeding the cap as a signal about the submission rather than an error in your pipeline. The same applies to the result document: a test framework can be made to emit an enormous report, and that report crosses a trust boundary on its way to your database.
What "integrity signal" means and does not mean
The product produces an integrity score. It is worth being precise about what that can honestly be.
It can measure things that are observable: how the work arrived, whether the session behaved consistently, whether the submission pattern matches the timeline. Those are signals, they have base rates, and they are useful to a human reviewer.
It cannot determine whether someone cheated. The cost of a false positive falls entirely on a candidate who may lose an opportunity over a statistical artifact, and that asymmetry means the product surfaces signals and never renders a verdict. A number with an explanation attached, handed to a person who decides.
That is a product constraint before it is a technical one, and it changes the engineering: every signal has to be explainable, which rules out anything whose output cannot be traced back to something concrete you could show the candidate.
What I would do differently
The one thing I underestimated was image lifecycle. Fourteen language images, each with a runtime and a test framework, each needing security updates, is an ongoing operational commitment that does not show up in the initial design and does not stop. If I were starting again I would build the rebuild-and-verify pipeline in week one rather than week twenty, because retrofitting it around images already in production is meaningfully harder.
The other half of running this alone is that the resource limits above are also the cost controls — the same reasoning that appears in the free-tier quota that took my product down for two days, from the other direction.