Forward-deployed engineer interview guide
What a forward-deployed engineering interview actually assesses, stage by stage — and how to prepare for each part without guessing.
Forward-deployed engineering is a delivery role. You are placed with a customer, you inherit their systems, and you are expected to leave behind something that keeps working. Interviews for the role are built around that, which is why they look less like a coding exam and more like a series of conversations about judgement.
Two things worth saying plainly. First, there is no single standard interview loop for this role — companies differ in length, order and format, so treat the stages below as the parts you are likely to meet, not a fixed sequence. Second, this guide is our own synthesis of what the work involves; the only hard numbers on this page are the counts above and the skill frequencies at the bottom, both recomputed from the live postings on every build.
Before anything else: what the role is
An FDE sits between a product and a customer's reality. The work is scoping, integration, evaluation, and keeping something alive once the demo is over. Everything an interview asks you is downstream of that: can you build inside constraints you did not set, decide what not to build, find out why something broke in a system someone else designed, and stay credible in front of the person paying for it.
What the process is testing
Four things come up in some form in almost every loop. Read these as the four evaluation axes; the stages below are just the vehicles.
Shipping inside someone else’s constraints
What it tests. Whether you can deliver working software in an environment you did not choose — the customer’s cloud, their security review, their existing systems, their release windows.
How it shows up. Expect questions about a time the environment fought you: a bank that would not allow egress, a client whose data could not leave a region, a review that took six weeks.
Strong answers. You describe the constraint precisely, then the one thing you changed to work within it, then how you proved it still met the requirement. Constraints are treated as design inputs.
Where it goes wrong. Stories where you replaced the customer’s stack with your preferred one, or where the constraint is never mentioned until the end as an excuse.
Turning a vague problem into something buildable
What it tests. The job almost never arrives as a ticket. It arrives as “our teams spend too long on this and we think AI can help”.
How it shows up. A short case: a stakeholder describes a workflow badly. You are asked what you would do in the first two weeks.
Strong answers. You ask about volume, error cost, who reviews the output, what the current workaround is — then propose the smallest version that produces a measurable change.
Where it goes wrong. Jumping straight to architecture. Assuming the stated problem is the real problem. No mention of how anyone would know it worked.
Debugging systems and data you did not design
What it tests. Most of the failures you will be paged for are not model failures. They are a stale field, a duplicate key, a permissions change, a pipeline that silently started returning empty results.
How it shows up. A live debugging exercise, or a walkthrough of the worst production incident you have owned.
Strong answers. You separate “what I observed” from “what I concluded”, name the instrument you used, and say how you would have caught it earlier.
Where it goes wrong. A tidy narrative with no uncertainty in it. Anyone who has actually been on call has at least one hypothesis they were wrong about.
Being trusted in front of the customer
What it tests. FDEs sit with the customer. The company is lending you their relationship, so they are checking whether you are safe to send.
How it shows up. A role-play or a discussion of a disagreement with a stakeholder — including one where you were wrong or had to say no.
Strong answers. You can explain a technical trade-off in the language of the person paying for it, and you escalate disagreement without going around anyone.
Where it goes wrong. Contempt for non-technical stakeholders, or a story where every conflict was resolved by being right.
The stages you are likely to meet
Order and depth vary by company and by level. A senior loop tends to spend more time on scoping and design; a junior loop spends more on the applied round. Prepare for all of them — the marginal cost is small and the cost of being surprised is a rejection.
Recruiter or intake screen
Covers. Logistics, timeline, work authorization, compensation expectations, why this role.
What it is checking. Whether your expectations and the posting are the same job. A mismatch here ends the process cheaply for both sides.
How to prepare. Have your salary range ready and grounded — see the pay section below. Be able to say in two sentences why this role specifically, not “AI” in general.
Hiring manager conversation
Covers. Your background, how you work, what you have shipped end to end, what you want next.
What it is checking. Whether you have owned delivery rather than contributed to it, and whether the scope of the role matches your seniority.
How to prepare. Prepare two or three projects you can talk about in depth, with the trade-offs and the parts that did not work.
Technical screen
Covers. Programming, usually with data or strings rather than algorithmic puzzles; sometimes debugging an existing script.
What it is checking. Fluency and the ability to reason out loud. For this role family, working code beats clever code.
How to prepare. Practise in the language you would actually use on the job. Be comfortable reading unfamiliar code before you write any.
Applied or build round
Covers. A small build: a retrieval flow, an evaluation harness, an integration against a documented API. May be a take-home or a paired session.
What it is checking. How you scope when the brief is deliberately underspecified, and what you choose not to build.
How to prepare. Ask up front what “done” looks like and how long they expect it to take. Write the README before the code — they read it.
System and data design
Covers. A whiteboard or shared-document discussion of an end-to-end system: ingestion, retrieval, evaluation, deployment, cost.
What it is checking. Whether you design for the customer’s environment and constraints, and whether you can talk about failure modes and cost.
How to prepare. Bring your own defaults: how you handle permissions, what you log, how you evaluate before shipping, how the bill scales.
Customer-facing or scoping round
Covers. Often a role-play with a skeptical stakeholder, or a discussion of a deployment that went badly.
What it is checking. Whether you can be put in front of a paying customer without supervision.
How to prepare. Practise explaining one technical decision three ways: to an engineer, to a product owner, and to a finance lead.
Cross-team or values round
Covers. Conversations with people you would work alongside — support, sales, product, or another FDE.
What it is checking. Collaboration under pressure, how you give and take feedback, how you behave when the customer is frustrated with something you did not cause.
How to prepare. Ask each person what a bad week looks like in this team. The answer tells you more than the job description.
References and offer
Covers. Reference calls, sometimes with a customer or a former stakeholder; then compensation.
What it is checking. Whether the people who worked with you would do it again.
How to prepare. Tell your references which project you would like them to speak to. If pay is discussed, anchor on published ranges rather than a single number.
Question bank
30 questions of the kind this role attracts, grouped by what they are probing. Use them for self-testing: answer out loud, in under two minutes, and stop when you notice you are describing technology instead of a decision.
Scoping and ambiguity
- Tell me about a time a customer asked for something that turned out to be the wrong problem.
What they are listening for: Whether you push back early with evidence, or build the wrong thing well. - You have two weeks before a demo. What do you cut first?
What they are listening for: Whether you can rank by what the customer will actually judge, not by what is technically interesting. - How do you decide something is ready for a real user?
What they are listening for: Whether you have a definition of done that includes evidence, not just “the code works for me”. - A stakeholder wants a feature you believe will not survive contact with their data. What do you do?
What they are listening for: Whether you disagree with data and without contempt. - How do you write down scope so it does not drift?
What they are listening for: Whether you have a working habit — a written v1, explicit non-goals — or you improvise each time. - What is the smallest version of this that would change behaviour?
What they are listening for: Whether you default to shrinking the problem rather than expanding the architecture.
Building and shipping
- Walk me through something you shipped from first conversation to production.
What they are listening for: Whether you owned the whole path, including the parts that were not engineering. - How do you handle a security review that blocks your access?
What they are listening for: Whether you treat the customer’s process as a constraint to design around, rather than an obstacle to complain about. - What do you do when the model is good enough but the product is not?
What they are listening for: Whether you know the difference between model quality and the surrounding workflow — most deployment failures live in the second one. - How do you roll something back at a customer site?
What they are listening for: Whether rollback was designed in, or discovered during an incident. - Describe your default logging and monitoring for a new service.
What they are listening for: Whether you instrument before you need to, and what you consider a signal versus noise. - When do you choose not to use a model at all?
What they are listening for: Whether “AI” is a solution you reach for or the field you insist on. The best answer usually includes a rule or a lookup table.
Debugging and data
- A customer says the assistant “is getting worse this week”. How do you investigate?
What they are listening for: Whether you reach for evidence — logs, eval results, a changed field — before you reach for a prompt rewrite. - Tell me about the worst production incident you owned.
What they are listening for: Whether the story contains a moment where you were wrong, and what changed afterwards. - How do you tell whether a retrieval problem is actually an ingestion problem?
What they are listening for: Whether you trace the pipeline backwards instead of tuning the part that is easiest to change. - What does a silent data failure look like in your stack?
What they are listening for: Whether you know that the expensive bugs return empty results successfully rather than raising errors. - How do you test a system whose output is non-deterministic?
What they are listening for: Whether you have an evaluation set and a tolerance, or you test by reading a few outputs and hoping. - A field you depend on is renamed upstream. How do you find out before the customer does?
What they are listening for: Whether your monitoring is on the contract with the data, not only on your own service.
Customer and communication
- Explain embeddings to a customer who has never heard the word.
What they are listening for: Whether you can calibrate depth to the listener instead of to your own comfort. - Tell me about a time you had to say no to a paying customer.
What they are listening for: Whether you can hold a technical line without damaging the relationship. - A customer asks for a guarantee you cannot give. What do you say?
What they are listening for: Whether you are honest about uncertainty — this is the single most common way a deployment sours. - How do you run a weekly status update for a non-technical sponsor?
What they are listening for: Whether you report progress and risk, or only progress. - Someone on your team tells you your design is wrong in front of the customer. What do you do?
What they are listening for: Whether you keep the disagreement professional and move it off-stage. - How do you hand a system over so it keeps running without you?
What they are listening for: Whether you build for the customer’s team, including the documentation and the runbook.
Measurement, evaluation and cost
- How did you know your last deployment was working?
What they are listening for: Whether “it works” means a number, or means a feeling. - What would you measure in the first week at a new customer?
What they are listening for: Whether you instrument usage and quality before optimising anything. - How do you build an evaluation set when there is no labelled data?
What they are listening for: Whether you know how to start small and honestly, and how you would grow it. - Where does the cost of this system actually accumulate?
What they are listening for: Whether you understand tokens, retries, context size, and that a retry loop is a billing event. - How do you decide between a bigger model and a better pipeline?
What they are listening for: Whether you reach for capability before you have removed the obvious waste. - What would make you tell a customer to stop this project?
What they are listening for: Whether you can imagine a negative result, which is the difference between an engineer and a vendor.
Questions worth asking them
An interview is a two-way evaluation, and the questions you ask are part of the assessment. These are ordered by who can usually answer them.
- Hiring manager: What did the last person in this role ship that you were happy with — and what would you have done differently?
- Hiring manager: How many engagements is an FDE on at once, and who decides when one is over?
- Hiring manager: What percentage of the work is on-site or with the customer, and how is travel handled in practice?
- The team: What does a bad week look like here?
- The team: Who owns the customer relationship when something breaks at 2am — and how does that get handed over?
- The team: How much of your time goes to new deployments versus keeping existing ones healthy?
- Engineering leadership: How do you decide whether a deployment failed because of the model, the data, or the process?
- Engineering leadership: Where does the evaluation budget sit — is building eval sets part of the job, or someone else’s?
- Recruiter: What does the compensation band for this level look like, and how is it split?
- Recruiter: What has made recent hires in this role succeed or leave in the first year?
A preparation sequence
Not a timetable — a dependency order. Each step makes the next one cheaper.
- 1. Pick three projects and write them down. One where you shipped something end to end, one where you debugged something you did not build, one where you disagreed with a customer and it went well. Write each as: the constraint, what you did, the evidence it worked, what you would change. Most interview answers fail because the story is only in the candidate’s head.
- 2. Read the posting properly, then check what the company actually builds. The posting tells you the delivery surface — the industry, the systems you will integrate with, whether the work is on-site. The company’s own product pages tell you what a deployment looks like there. Both are useful; neither is a promise about the interview.
- 3. Do one applied exercise out loud. Take any of the questions below and build the smallest working version while narrating. The skill being tested in a build round is scoping and communication as much as code, and narration is the part nobody practises.
- 4. Prepare your defaults. Have a standing answer for: how you evaluate before shipping, how you handle permissions and data residency, what you log, and how cost scales. These come up in almost every design conversation and are much stronger as habits than as opinions invented on the spot.
- 5. Decide your compensation position before the first call. A range you can defend is better than a number you invented. The pay section below shows what employers on this site have actually published, with the sample size stated. Use it as a floor for the conversation, not as a ceiling.
Where candidates lose the offer
- Answering the interview you expected instead of the one you are in. If the round is about a customer conversation, do not steer it back to architecture.
- Telling the story without the constraint. “We built a RAG pipeline” says nothing; “the data could not leave the region, so we did X” says everything.
- Claiming a number you cannot source. One invented metric in an interview answer costs more trust than a plain “we did not measure that”.
- Treating the take-home as a coding exam. It is a scoping exam with code attached — an over-built submission usually reads worse than a small, finished one with a clear README.
- No questions. An interview is a two-way evaluation, and candidates who ask nothing read as either uninterested or unprepared.
- Silence while thinking. For this role the process is the product; narrating your reasoning is part of what is being assessed.
What the live postings ask for
The table below counts the skills employers themselves wrote into their postings for these roles. It is a picture of the delivery surface, not a syllabus: a skill appearing in a posting means the employer listed it, not that any particular interview will test it. Each skill links to the live roles that mention it.
- Agents 220 live roles
- Python 137 live roles
- AWS 96 live roles
- GCP 90 live roles
- Azure 76 live roles
- Observability 65 live roles
- SQL 64 live roles
- Evals 54 live roles
- Kubernetes 51 live roles
- Anthropic API 50 live roles
- RAG 50 live roles
- APIs 43 live roles
Counts are recomputed from the live job set on every build and cover only skills in the site's curated skill list — currently 36 skills with at least one match, of which 35 also have their own browse page. Free-text skill strings that do not map to a curated entry are not counted here.
Pay
Compensation comes up in the first call far more often than candidates expect. Do not invent a number: the employer has to justify theirs, and you should be able to justify yours. We publish the ranges employers actually wrote into their postings, with the sample size stated, on the salary data page — employer-disclosed figures only, nothing estimated.
One last thing
If you take nothing else from this page: narrate your reasoning, name the constraint, and be explicit about what you would need to find out. That is the behaviour the role rewards, and it is the behaviour the interview is built to detect.
Role counts and skill frequencies are recomputed from this site's live job set on every build (currently 440 open roles across 154 companies). The interview guidance itself is written by this site from the shape of the work — it contains no survey data, no pass-rate statistics, and no claims about any named company's process.