Machine codingEdtech at scaleDoubt queues & tests

PhysicsWallah-Style Machine Coding Round: Format, Tips & Practice Problems

Written and reviewed by Sahil Srivastav

A mass-market edtech platform has a different engineering shape from a premium one: enormous cohorts, content organised into batches and chapters, and support systems — doubt solving, test series, ranks — that must stay responsive when a million learners open the app after school. Machine coding rounds in the style of PhysicsWallah reflect that, with scenarios about routing, grading, and ranking rather than one-to-one scheduling.

The doubt queue is the most characteristic of them: learners post questions, they route to teaching assistants by subject and language, they escalate when unanswered, and no two assistants may pick up the same doubt. This page covers the reported format, what evaluators weight, and Gronex repositories that train the same routing and grading mechanics.

What a PhysicsWallah-style machine coding round looks like

Typical statements are a doubt-routing service (queue doubts by subject, assign to eligible assistants respecting load limits, escalate after an SLA, reassign on abandonment) or a test-series grading engine (submit an attempt, grade objective and multi-correct items under a negative-marking scheme, publish a rank list). Both come with precise rules and an obvious source of ambiguity the round wants you to resolve explicitly.

The doubt queue is an exactly-once assignment problem with an SLA on top. Two assistants polling simultaneously must not receive the same doubt, an assistant who goes idle must have their doubt returned to the queue, and escalation must fire from an injected clock rather than a live timer. This is the same claim-with-lease pattern as job dispatch, and rounds in this style expect you to name it.

Grading is graded on exactness. Negative marking, partial credit for multi-correct answers, and unattempted-versus-wrong are three distinct cases, and the statement’s scheme must be implemented literally. Rank lists then need a deterministic tie-break — usually fewer wrong answers, then earlier submission — and the reviewer will construct an exact tie to check that two learners never share a rank ambiguously.

How you’re evaluated

Exactly-once doubt assignment

Concurrent pickup never duplicates a doubt, abandonment returns it once, and a lease or claim makes that automatic.

Eligibility and load limits

Routing by subject and language with per-assistant concurrent limits, and deterministic selection among eligible candidates.

Clock-driven escalation

SLA breaches evaluated against an injected clock so escalation is testable and reproducible.

Exact grading and ranking

Negative marking, partial credit, and unattempted handled as separate cases, with a fully specified rank tie-break.

Common mistakes that fail this round

  • Assigning doubts with a read-then-mark pattern, so two assistants answer the same question.
  • Returning abandoned doubts only on explicit release, losing them when an assistant disconnects.
  • Using real timers for escalation, which cannot be tested inside the round.
  • Treating unattempted answers as wrong, silently applying negative marks the scheme does not allow.
  • Producing a rank list whose order changes between runs because the tie-break was left implicit.

Quick tips for the room

  • Claim with an owner and a lease; a busy flag cannot survive a disconnect.
  • Write one comparator for assistant selection and one for rank ordering.
  • Handle unattempted, wrong, and partially correct as three distinct branches.
  • Drive every SLA and escalation from an injected clock.

How to prepare

Implement claim-with-lease for the doubt queue and prove with a test that two claimants get different doubts and that an expired lease requeues one. Then write the grading function with a case for each answer state and a comparator for ranks with every tie-break named. Those two artefacts are the round in miniature.

The repositories below cover them: the CRM assignment problem is capability routing with escalation, the distributed scheduler is exactly-once claiming with leases, the SLA problem is clock-driven breach evaluation, the exam grading engine is the scoring and ranking core, and the free job-queue extension is deterministic dispatch with priority.

Practice problems in the PhysicsWallah-style round format

Each is a real backend repository with a failing test suite — the same working-code standard the round applies. Open the brief and read the full problem, no signup required.

MEDIUM~90 min

Assignment & Escalation System

Routing by capability with load limits, deterministic selection, and escalation on timeout.

Open the challenge →
HARD~120 min

Distributed Job Scheduler

Exactly-once claiming with leases when many workers race the same item — the doubt-queue core.

Open the challenge →
MEDIUM~75 min

Attempt & Grading Engine

Grading with negative marking and partial credit, attempt state, and deterministic result publication.

Open the challenge →
MEDIUM~90 min

SLA & Escalation Tracking

First-response and resolution clocks with an injected clock and testable breach evaluation.

Open the challenge →
MEDIUMFree~45 min

Job Queue: Delayed & Priority

Deterministic dispatch: due-time gating, priority ordering, FIFO tie-breaks. Free to try.

Open the challenge →

Rehearse the round before you sit it

Open a real repository, see the failing tests, and make them pass against the clock — the loop a PhysicsWallah-style machine coding round actually grades. Start free, no card required.

FAQ

Are these actual PhysicsWallah interview questions?

No. They are Gronex originals in the style of large-scale edtech rounds — the kind of problem asked in rounds like PhysicsWallah’s. Gronex is not affiliated with PhysicsWallah.

Why is the doubt queue a good interview problem?

Because it combines three things in one small scope: exactly-once assignment under concurrency, eligibility-based routing, and a clock-driven SLA. Each is individually testable, which is why reviewers like it.

How exact does grading need to be?

Literal. If the scheme says minus one for wrong and zero for unattempted with partial credit on multi-correct, all three must be implemented separately. Reviewers hand-compute an expected score and compare.

Do I need to think about scale in the build?

Not in code, but expect the question. Good answers name the hot spots — the queue, the rank list — and describe partitioning by subject or batch, with the rank list computed periodically rather than on every submission.

Related

Gronex is not affiliated with, endorsed by, or sponsored by PhysicsWallah. All company names and trademarks belong to their respective owners. The problems on this page are Gronex originals written in the style of such interview rounds — not actual interview questions from PhysicsWallah.