skip to content

Features

Real cluster. Real fault. Real engineer. Find out.

The mechanics in detail: how the environments are built, what the evaluator watches, how a session is graded, and what we deliberately refuse to build.

5
Stacks at launch
<5s
Environment boot
100%
Of every session
0
Cameras, ever

01

Real infrastructure, the whole DevOps surface

Live Kubernetes, Docker, Linux, Terraform, and a real monitoring stack (Prometheus, Grafana), a genuine environment with a genuine fault injected, booted in under 5 seconds. This is not a Kubernetes-only tool. DevOps means the pod that will not start, the container image that lies, the Terraform state that drifted, the dashboard that says everything is fine while everything is on fire. Code sandbox vendors cannot provision this; that is an architecture gap, not a feature gap.

Lab library120+ labs
allkubernetesdockerlinuxterraformmonitoring
  • Broken etcd cluster
    #kubernetes
    hard· 45m
  • Image that lies
    #docker
    medium· 30m
  • Disk full, inodes gone
    #linux
    medium· 30m
  • Terraform state drift
    #terraform
    hard· 45m
  • The lying dashboard
    #monitoring
    medium· 30m

02

Pick labs yourself, or upload the JD and let us recommend

Browse 120+ ready-made labs across every stack we support, filter by tool and difficulty, and send invites. Or skip the browsing: upload the job description and let our AI engine read the required skills and suggest the labs that actually test them. No DevOps background needed on your side. Add preferences if you like (must-cover tools, number of labs, time box per lab) and the suggestions respect them. Review, adjust, send.

Suggested labsJD: senior-platform-engineer.pdf
  • CrashLoop in payments
    #kubernetes
    96% match
  • Terraform state drift
    #terraform
    91% match
  • The lying dashboard
    #monitoring
    87% match
Your preferences: must cover Kubernetes + Terraform · 3 labs · 45 min each

03

AI evaluator: monitors everything, grades everything

An AI evaluator observes the entire session as it happens: every command typed, every exit code, every approach taken, every dead end entered and escaped. It generates the final graded report (diagnosis quality, time-to-first-correct-hypothesis, recovery from wrong turns, blast radius, AI usage), anchored by a deterministic check that verifies the system actually recovered. Nothing is sampled, nothing is skimmed: the whole session is the input.

Session report · #4-218PASS · strong diagnosis
  • Time to first correct hypothesis4m 12s
  • Hypotheses before correct2
  • Dead ends entered / recovered1 / 1
  • Commands that changed state6
AI evaluator watched 100% of the session, 214 commands

04

Blast-radius score

Deleted a namespace to make the symptom go away? Disabled a liveness probe and never re-enabled it? We saw. Nothing else on the market scores harm done while fixing. Recovering from your own mistakes counts for you, hiding them does not.

Blast radius1 finding
Disabled liveness probe, never re-enabled
at 09:14 · deployment/payments-api · still off at session end
No destructive shortcuts
nothing deleted to hide the symptom
Verdict: fixed the fault, left a landmine

05

AI-allowed by design

Every competitor sells AI detection, an arms race they lose. We grade how well candidates drive it instead: did they paste the right logs, catch the wrong suggestion, verify before applying? Honesty note: this is incentive design, not enforcement, a candidate can still use their own laptop, and we say so plainly.

AI usageeffective
  • Pasted the right logs on first ask
  • Rejected 1 incorrect suggestion (kubectl delete ns)
  • Verified the fix before declaring victory
Graded from the transcript, no detection, no proctoring

06

Session replay, asciicast, not screen video

Hiring managers do not trust assessment scores. They trust watching the fix happen. Every report links to a scrubbable replay of the full terminal session, every command, exit code, and dead end. And it is a terminal recording, not screen video: kilobytes instead of gigabytes, searchable and diffable text, and nothing from the candidate’s screen, camera, or mic ever leaves their machine. 90 seconds of scrubbing beats an hour of second-guessing a number.

Replay · session #4-21741 KB of text
[09:14:02] $ kubectl edit deploy payments-api
[09:14:31] $ kubectl delete probe... ^C # ← the moment
[09:15:07] $ kubectl rollout status deploy/payments-api
grep the session
kjstep commands
command 84 / 214
No video. No camera. No mic. Ever.

07

Randomized variants

One scenario, many faults. The broken-etcd lab has a dozen distinct root causes; the invite picks one at random. Leaked walkthroughs are worthless the moment they are posted, and your loop stays fair without a question-bank treadmill.

broken-etcd · variants12 variants
#01 wrong startup flags#02 corrupted member#03 expired peer certs#04 full data dir#05 clock skew#06 revoked lease+6 more
Invite picks one at random, leaked walkthroughs go stale instantly

Roadmap

What's next

Not shipped yet, listed anyway, beta feedback decides the order. What is deliberately not here: screen recording, webcams, or proctoring of any kind. The replay is the terminal, full stop.

In-session AI, prompt trail graded

A frontier-model assistant inside the session terminal. The full prompt trail is graded alongside the commands: what context they gave it, which suggestions they caught as wrong, whether they verified before applying. Bring-your-own-AI stays allowed either way.

Live shadow mode

Watch the session in real time, read-only, and the candidate always knows. For panel interviews that want to talk through the incident afterward, not surveillance.

Team incident drills

Same scenarios, aimed at the team you already have: scheduled game days, blast-radius scores per engineer, and a report you can bring to the postmortem. Practice the 3 AM page at 3 PM.

More stacks

Postgres on fire, BGP misbehaving, a broken CI/CD pipeline, a cloud bill exploding. Kubernetes, Docker, Linux, Terraform, and monitoring at launch; databases, networking, CI/CD, and cloud-provider scenarios next.

Candidate skill profile

A shareable, verified record of passed scenarios: what was broken, what they fixed, and what the fix cost. Signal that travels with the engineer, not the employer.

Debrief pack for the next round

The report already knows where the interesting moments were. Turn them into a short list of questions for the human round: ask about the probe they disabled at 14:02, the theory they abandoned, the check they skipped. We do not tell you who to hire, we make the conversation you have next a better one.

Environments, data, and what we refuse

The questions your security reviewer will ask, answered before they ask them.

Disposable per session

Every session gets its own environment, built for that attempt and destroyed when it ends. Candidates are expected to break things, so nothing they touch is shared with anyone else.

Nothing to install

The candidate needs a browser. No VPN, no cloud account, no agent on their machine, no admin rights. Nothing from their laptop is touched.

What is recorded

Terminal commands, their exit codes, and diffs of files changed inside the environment. That is the complete list, and it is what the replay is made of.

What is never recorded

No camera, no microphone, no screen capture, no biometrics, no keystroke dynamics. None of it exists in the product, and adding any would mean changing the privacy policy first.

Retention and deletion

Sessions and reports are kept for 12 months and then deleted. You can delete a candidate session earlier at any time, and candidates can request deletion of their own data.

Your code stays yours

Labs run on our scenarios, not your repositories. A custom lab is built to resemble your stack, not to contain it, so nothing proprietary leaves your side.

Full detail in the privacy policy.

How Faultybox compares

Categories, not names, you know who they are.

FaultyboxCode-sandbox platformsQuiz platforms
Real infrastructurelive k8s / docker / linux / terraform / prometheussandboxed snippetsmultiple choice
AI evaluationfull session monitored + gradedoutput onlyanswer only
Test selectionJD-matched lab suggestionsmanual pickingfixed question banks
Blast radiusscored
AI policyallowed + gradeddetection arms racebanned / proctored
Replayevery sessionsometimes

The box is faulty. Want in?

Private beta, limited seats. Engineers get free sessions; hiring teams get white-glove setup.