← Live feed

Distributed Systems Engineer — Diagnose a Non-Deterministic Production State Failure

Budget: $5550.0 FIXED / ⭐ 0.00 (0) ARE

next.js, typescript, react-js, software-debugging, postgresql, performance-testing, devops

Preferred qualifications

  • Experience: Expert
We are looking for an exceptionally strong Staff or Principal level engineer for a difficult systems reasoning and reliability engagement involving a modern distributed web platform. This is not a standard frontend development role. The first stage is a technical reasoning screen built around a synthetic production incident designed to evaluate how an engineer handles contradictory evidence, distributed state, nondeterminism, caching, rendering boundaries, concurrency, observability, and test isolation. We are specifically interested in engineers who can determine when the evidence model itself is wrong rather than immediately proposing framework level fixes. SYSTEM CONTEXT Assume a production system with the following architecture: Next.js App Router React Server Components Client Components Streaming server side rendering Partial rendering Client side navigation Route prefetching TypeScript Independent API service PostgreSQL primary and read replicas Redis CDN and edge caching Background revalidation Localized routing Authenticated and public request contexts URL driven application state Optimistic client transitions Distributed tracing Browser automation Unit, integration, contract, and end to end test layers For this scenario, request identity is intended to be a deterministic function of: tenant deployment revision locale authentication scope pathname normalized query state relevant feature configuration The expected invariant is: Given equivalent authoritative inputs and the same committed application revision, the externally observable result should converge to the same logical state regardless of whether the route is reached through a direct request, an RSC transition, prefetch plus navigation, browser history restoration, server rendering, or client reconciliation. THE INCIDENT An intermittent production incident has produced the following observations. All observations originate from different pieces of instrumentation and must not be assumed equally trustworthy. A Two users request what the system believes to be the same logical route. Both requests have identical deployment revision, normalized URL, tenant, locale, authentication scope, and feature configuration. The RSC payload captured for both requests is reported as byte for byte identical. The JavaScript bundles are also reported as byte for byte identical. However, after hydration completes, the final DOM trees differ. No hydration warning is recorded. No React state update is recorded after hydration. The difference persists when the interaction is repeated. B Both browser environments are reported to have identical browser version, viewport, timezone, locale, reduced motion preference, device pixel ratio, no extensions, no service worker, no local storage, no session storage, no IndexedDB, no persisted HTTP cache, and no cookies other than the explicitly captured authentication state. No use of wall clock reads, application randomness, cryptographic randomness, browser specific feature branching, or third party scripts is visible in the application render path. Despite this, reversing which machine receives which captured network session reportedly preserves the divergence by machine rather than by network capture. C The API response independently captured for the affected route is correct. The API service reports that the request reached PostgreSQL. PostgreSQL logs show one query corresponding to the request. Distributed tracing shows two database spans. Redis reports neither a cache hit nor a cache miss. The CDN reports a cache hit. The application server simultaneously reports that it generated the response body. A single request identifier appears across all of these records. D The issue reproduces more frequently after client side navigation than after direct navigation. Refreshing the page commonly removes the symptom. Back and forward navigation can reintroduce an earlier visible state even though the URL remains correct. Disabling route prefetching significantly reduces reproduction frequency but does not eliminate the issue. Disabling application caching also reduces reproduction frequency. Neither action eliminates the underlying class of failure. E A second symptom appears only under an RTL route. The visible failure looks almost identical to the original incident, but occurs through a different sequence of user interaction. There is no evidence yet that the two symptoms share a root cause. F The automated test suite currently reports 37 failures. Each failed test passes when executed individually. Every pair of previously failing tests also passes when executed together. Any selected group of three tests produces at least one failure. The identity of the failing test changes with execution order. The test infrastructure claims that every test receives an isolated browser context, a unique database schema, a unique Redis namespace, frozen application clocks, deterministic random seeds, independent authentication state, and no shared filesystem state. Unexpectedly, increasing worker count from 1 to 8 reduces the observed failure rate. Adding a one millisecond artificial delay to an unrelated assertion causes all 37 failures to disappear. Increasing normal test timeouts does not produce the same result. G The following remediation attempts are not acceptable unless the engineer first proves they address the root cause: Disabling server side rendering Disabling React Server Components Disabling prefetching globally Disabling caching globally Converting all routes to dynamic rendering Forcing browser reloads Moving all authoritative state into the client Increasing arbitrary sleeps or timeouts Weakening assertions Serializing the entire test suite Introducing another state management framework Rewriting the application YOUR FIRST TASK Do not propose a fix. Do not write code. Do not give us a generic debugging checklist. Before suggesting any remediation, analyze the evidence model itself. We want you to answer the following: 1. Which observations above cannot safely be accepted as independent facts? 2. Is there any subset of the reported observations that appears logically incompatible under the stated execution model? 3. What is the minimum set of assumptions that must be false, incomplete, or incorrectly measured for the incident to remain physically possible? 4. Rank the available evidence from highest to lowest diagnostic reliability. Explain why. 5. Give your first five investigative actions in order. For each action provide: The hypothesis being tested The exact evidence you want The expected outcomes What would falsify your current hypothesis What the next branch of investigation becomes under each outcome 6. Identify at least three places where two apparently independent observations could actually originate from the same flawed instrumentation or correlation boundary. 7. Explain how you would distinguish among incorrect server state, stale RSC state, incorrect cache identity, stale client state, history restoration, hydration or reconciliation behavior, request correlation failure, hidden test contamination, and scheduler or timing sensitivity without immediately changing production behavior. 8. The CDN reports a cache hit while the origin claims to have generated the response. Give multiple explanations under which both statements could be operationally true. Then explain what evidence would distinguish them. 9. PostgreSQL reports one query while the trace reports two database spans. Do not assume either side is wrong. Describe several models that could produce this observation. 10. Why is the following statement dangerous? The API response is correct, therefore the backend is not responsible. 11. What would make the one millisecond delay observation useful evidence rather than merely a coincidence? 12. Why might increasing concurrency reduce the reproduction rate of a race, ordering, isolation, or contamination bug? 13. Identify one part of this incident that you would deliberately refuse to investigate during the first phase. Explain why investigating it early would reduce information gain. 14. Finally, state your current leading model. Then state the single observation that, if produced tomorrow, would most strongly force you to abandon that model. WHAT WE ARE EVALUATING We are not grading candidates on whether they guess a hidden answer. We are evaluating: Ability to reason from invariants Evidence ranking Falsification discipline Causal reasoning Understanding of distributed state Understanding of browser and server boundaries Ability to detect impossible premises Ability to distinguish correlation from causation Test isolation reasoning Observability skepticism Ability to reduce a large incident into minimal discriminating experiments Willingness to say that the available evidence does not support a conclusion The strongest candidate may conclude that parts of the incident report cannot simultaneously be true. That is acceptable. In fact, discovering that an observation is impossible under the stated model may be more valuable than producing a sophisticated explanation for it. WHAT WE DO NOT WANT Please do not send generic framework debugging advice. Please do not send a list of libraries you would install. Please do not send a rewritten architecture. Please do not send a long explanation of what React hydration is. Please do not suggest clearing Redis, disabling cache, adding retries, increasing timeouts, introducing Redux, using useEffect, or turning off prefetch without a falsifiable diagnostic model. APPLICATION Your proposal should contain your response to the reasoning challenge above. Concise answers are welcome. A strong two page analysis is substantially more valuable to us than ten pages of generic material. Candidates who demonstrate exceptional diagnostic reasoning may proceed to a deeper technical discussion and a larger fixed price engineering engagement. No repository access is required for the initial stage. We are looking for engineers who are comfortable being challenged on their reasoning and equally comfortable challenging our assumptions when the evidence does not support them.
Open job

AI proposal draft

Generate a short cover letter to copy into the offer. Says you are interested and ready to work.

Sign in to generate an AI proposal draft.

Log in