skip to content
‹ All posts

AI-allowed interviews: stop policing, start measuring

Everyone is fighting AI cheating in interviews with detection tools. We think the arms race is unwinnable - and the premise is wrong. Measure instead.

#opinion

The hottest product category in hiring right now is surveillance.

Search for "AI cheating in interviews" and you'll find an industry in a panic spiral: detection vendors, browser lockdowns, eye-tracking, second-camera requirements, think-pieces about candidates with hidden earpieces. The entire conversation assumes one framing - AI use is contamination, and the job of an assessment is to keep the sample sterile.

We take the opposite position, and we're not being contrarian for sport. AI-allowed is the only framing that survives contact with reality. Here's the argument, in three parts.

The "AI cheating in interviews" framing loses on its own terms

Suppose you accept the premise that AI use is cheating. You still lose, because detection is a structurally losing game.

Detection models trail generation models. This isn't a temporary gap someone will close with a funding round; it's the shape of the problem. The generator only has to produce output a classifier can't distinguish from human work, and every improvement in classifiers is training signal for generators. You are betting your hiring pipeline on permanently winning a race in which your opponent studies your last move before making theirs.

Meanwhile the costs land on the wrong people. Every false positive is a real candidate - likely a good one, since fluent output is what triggers suspicion - accused of fraud by a probability score. Every lockdown requirement tells senior engineers, the people with the most options, that your process starts from distrust. The proctoring arms race is a tax you pay to make your funnel worse.

The premise is wrong anyway

Now reject the premise, because it deserves rejecting. Your engineers use AI at work. Today, on real tickets, with your blessing - most engineering orgs are actively pushing adoption. An interview that bans AI is a simulation of a job that no longer exists. You are screening candidates for their performance in a counterfactual world, and then surprised when interview performance doesn't predict job performance.

The take-home already died this way - we wrote its obituary. The artifact-as-proxy model broke the moment artifacts became free. The answer isn't to guard the proxy harder. It's to stop measuring proxies.

What to measure instead

Here's what changes when you put a candidate in front of a genuinely broken system - a live cluster with a real fault - and tell them AI is allowed:

Cheating becomes structurally impossible, because there's nothing to cheat with. The skill under test includes the model. Pasting the error into an LLM isn't a violation; it's move one, same as it would be on-call. What you watch is everything after:

  • Did they give the model the right context, or paste a stack trace with the actual signal cropped out?
  • When the model suggested something plausible and wrong - it will - did they catch it, or apply it blindly and compound the outage?
  • Did they verify the fix themselves, or accept "that should work" from a system that cannot see their cluster?
  • Were they driving, or being driven?

That last distinction is the whole game. An engineer steering an LLM through an incident is displaying a real, increasingly central job skill. An engineer being dragged by one is displaying the most expensive failure mode in modern operations: confident automation of a wrong idea. A transcript shows you which one you're watching, timestamp by timestamp. No detector required - the evidence isn't whether AI appeared, but what the human did with it. Candidates preparing for this kind of interview can read how to use AI well when it's allowed; the short version is that the model is a power tool, and we're grading your grip.

This is why Faultybox sells no AI detection and no proctoring. Not as a feature gap - as a position. Deterministic scripts check whether the system actually recovered. The recorded session shows how. Nowhere in that pipeline does anyone need to guess what species wrote a given command.

"But then what stops them from just letting the AI do it?"

Nothing. That's the point. If a candidate can hand a live, faulted, unfamiliar production environment to an AI, supervise it to a verified fix inside 45 minutes without breaking anything else - congratulations, you've found someone who can do that, which is rapidly becoming the job description. In practice, today's models let go of the wheel at exactly the moments that matter: stale assumptions, misleading symptoms, fixes that need verification against ground truth. The candidates who shine are the ones who know when to grab it back.

The detection industry needs AI use to be cheating, because that's the product. You don't. Reframe the interview around the job as it exists, and the entire cheating question dissolves. Stop policing. Start measuring.


Faultybox runs your candidates through a real broken cluster and shows you exactly how they fixed it - replay included. Free pilot in beta → join