Real repositories three modes graded against the real fix
Production is two hundred files you did not write, a deploy that fails for reasons the error does not name, and logs as the only clue. Proddojo drops you into that, grades your fix, and shows the commit that fixed it for real.
$ docker logs checkout | tail -4 14:23:01 INFO Publishing order.created 14:23:01 WARN Queue connection timeout, retrying 14:23:07 WARN Queue connection timeout, giving up 14:23:07 INFO POST /checkout completed 200
a1c9e02
"fail the request when the publish gives up"
Readings from the last run.
The gap
Debugging under uncertainty is the skill that closes the distance, and it is the one nothing on the market teaches on purpose.
Proddojo is built entirely out of what production requires.
A self-taught developer or a bootcamp graduate can build an application from a clean start. The first production codebase is nothing like that: unfamiliar structure, a build that fails for reasons the error does not name, three services talking to each other, and a bug report that says orders are missing and nothing else.
Interview platforms test algorithms. Courses test syntax. What hiring managers test, in take-home tasks and first weeks, is whether someone can read logs, follow a stack trace through code they did not write, and change it without breaking the rest.
The bugs already exist. Open-source projects fix thousands of them a year, each with the report, the discussion and the commit that closed it. Packaged as challenges, they are the closest thing to a first week that can be practised.
The form
Four movements, in order, every time. The environment is live before the brief is read, and nothing is multiple choice.
A scenario in the words a customer or a colleague would use, and nothing else.
A real environment in the browser: the services running, docker logs, psql and git blame available. No multiple choice. Hints on request, from a model that has read the fix and will not reveal it.
The change runs against the project's own tests in an isolated container, with azkaban doing the isolation.
A second model grades the fix against a rubric built from the maintainer's actual commit, the way Mockterview grades an interview with a model that never conducted it. The real commit is shown beside yours.
Three modes
The same environment and the same grader, pointed at the three things a first month actually asks for.
A report that names a symptom. The cause is somewhere in a codebase nobody in the room wrote.
A change added to a repository that already has opinions, without breaking what depends on it.
A diff to read the way a colleague would, then approve or block it, with the reasons written down.
Every challenge in all three modes is drawn from an open-source project under a permissive licence, and the project, the issue and the fixing commit are credited on the challenge itself.
Progression
A rank moves when a challenge is solved without a hint and the grader agrees with the maintainer. It does not move for time served.
11 challenges, 2 with a hint
7 challenges, logs and traces
5 challenges, locks and migrations
9 challenges, contracts and retries
4 challenges, red pipelines
2 challenges, both with a hint
5 challenges, one profiler
6 diffs read, 1 blocked correctly
Designed progression, with example values.
Who it is for
Can build from a clean start and has never opened a codebase they did not write, with a deploy that fails and a report that names a symptom.
Wants the first month to be spent on the team's own codebase. Challenges built from the team's repositories, graded the same way, are the onboarding track.
The code arrives faster than anyone reads it, and the lead can no longer say who could debug the billing path at three in the morning. The team track keeps that answer current.
The team track
Short drills built from the team's own repositories and incident history, so the practice is on the systems each engineer is on call for, and it pays off at the next page rather than in a course certificate.
A team eighteen months into heavy agent use ships more code than it reads. Everything passes review, everything runs, and the skill that goes quietly is the one nobody measures: following a failure through code you did not write. It shows up as an incident that takes days, or a module that only one person could explain, and they have just handed in their notice.
Anthropic's own randomised study of junior engineers found that how they used the model mattered more than whether they did: the ones who asked it conceptual questions understood the code afterwards, and the ones who delegated the writing to it did not. That is a behaviour, and behaviours can be practised.
A realistic fault planted in a sandboxed copy of one of the team's own modules, by mutation rather than by hand, with the module's own tests as the judge.
Walk through a flow in your own words. The grader has read the actual code and checks each claim against it, not against a model answer.
A merged pull request, before its consequences. Say what fails in production and where, then see what actually did.
A past postmortem re-run as a timed exercise, with the logs and the repository as they were that day and the ending withheld.
An agent-written pull request with one real flaw in it, and every test green. Approve it or block it, and write down why.
Short enough to fit between tickets, and only worth doing if the team's lead has put the time in the week. Optional homework on a twelve-hour day does not happen.
A model is on hand throughout, and it answers with questions: where would you look, what does that log line rule out. It has read the answer and will not write it.
Each engineer sees their own results and nobody else does. The lead sees the team map below, which names modules, never people. A drill that feels like surveillance is a drill nobody takes honestly.
explain-back and bug hunt, 14 drills this month
one incident replay, median 41m to root cause
on the critical path; 80% of it merged from agents this year
no explain-back passed; review drills approved the planted flaw twice
Designed team map, with example values. A module at zero is the one to drill next, or to recover.
What it costs
Four tiers. Individuals pay per seat, monthly; teams pay per seat for challenges built from their own repositories.
free
Get your feet wet with real challenges.
per seat, monthly
Unlimited training across all three modes.
per seat, monthly
A model reading over your shoulder, for faster growth.
per seat, teams
Onboard faster, and keep the codebase understood after the agents moved in.
Request access
Requests decide which stacks the first challenges are drawn from.
Asked so far