What the job actually is
The title comes from forward-deployed military engineering: you are positioned at the edge of the organisation, close to the thing that matters. In Gen-AI that means inside a customer’s or a business unit’s environment rather than inside my own platform team. Concretely, the output of the role is a working system in someone else’s environment, plus the trust that lets it keep running:- I take a vague business complaint (“our support agents repeat themselves”, “the migration assessment takes three months”) and turn it into a measurable system with a latency and cost budget.
- I build the retrieval or agent pipeline end to end, including the parts nobody demos: evaluation, observability, guardrails, retries, and the invalidation path.
- I work next to people who do not read my code, so explanation is a deliverable, not a courtesy. This is why every FDE loop I have seen contains a teaching round rather than only a coding round.
- I own the trade-off conversation. The customer will get a model choice, a chunk size, and an SLA whether or not I chose them deliberately.
The skills the role assumes
The interview guide I have covers architecture, production systems, debugging, security, and engineering trade-offs, and the panel assesses understanding of Applied AI systems through scenario-based questions. The underlying competency map I would study:The interview loop
The loop I have documented is 50 to 60 minutes total, split into two parts.1
Part 1 — Demo teaching round, 30 minutes
I teach one topic — RAG, or AI agents and tool use — to a small group of new FDE hires who understand software engineering fundamentals but are new to Applied AI. I assume no prior knowledge and leave room for at least one audience interaction.
2
Part 2 — Knowledge and discussion round, 20 to 30 minutes
After the teaching session the panel asks scenario-based questions to evaluate my understanding of modern Applied AI systems: production challenges, debugging approaches, and architectural decision-making. Questions are scenario-based and designed to assess practical problem-solving rather than theoretical memorization.
The teaching round, decomposed
The expected structure is the rubric, so I build the session to match it exactly:- Introduce the real-world problem.
- Build intuition before introducing technical concepts.
- Explain the system using clear examples and diagrams.
- Discuss practical limitations and failure modes.
- Engage learners through questions or discussion.
- Summarize key takeaways.
- Chunking losing context, where the answer spans a boundary I chose for token-count reasons.
- Retrieval recall versus precision, and the fact that I can only generate over what survived top-k.
- Stale indexes, where the document changed and the vector did not.
- Lost in the Middle, where position inside a long prompt changes whether the model uses the evidence.
- Query-document mismatch, where the user’s words and the document’s words never overlap lexically.
- Runaway or looping tool calls, which turn one cheap request into an unbounded spend.
- Error propagation across multi-step workflows, where step 3’s bad output silently becomes step 7’s confident answer.
- Tool-call hallucinations, including invented parameter names and endpoints that do not exist.
- Ambiguous delegation between agents, where two agents both own the step, or neither does.
- Context or state loss during hand-offs, the multi-agent version of a dropped session.
The knowledge round
The topic list from the guide, with the shape of the answer I give for each:- Fine-tuning versus prompt engineering versus RAG. I frame it as what each changes: prompts change behaviour at zero data cost, RAG changes the facts available and keeps provenance, fine-tuning changes the weights and is the slowest and least auditable. Freshness and citation needs usually eliminate fine-tuning as the primary answer.
- Diagnosing production issues in RAG systems. I go stage by stage and measure: is retrieval missing the evidence, is the reranker reordering it away, is the context too long for the model to use it, or is generation ignoring grounded context? Each has a different metric and a different fix, and “the answers are bad” is not a diagnosis.
- Agent orchestration patterns. Single loop with tools, planner-executor separation, supervisor with worker agents, and workflow-first graphs where the LLM only fills in steps. I pick by how much autonomy the task actually needs and what a wrong step costs.
- Multi-agent system design. The interesting problems are shared state, hand-off contracts, and cancellation, not the org chart of agents. Most multi-agent designs I review would be better single-agent workflows with clearer tool schemas.
- Prompt injection and AI security. Retrieval and tool output are untrusted input. The model cannot separate data from instruction, so control has to live outside it: capability scoping, allow-listed tools, approval gates on writes, and least privilege on every credential a tool can reach.
- Guardrails and production best practices. Input filtering, output schema validation, PII redaction, confidence-based routing, retries with budgets, circuit breakers, and human approval on irreversible actions. See the reliability stack described in Sample Resume.
- Evaluation and observability of AI systems. Retrieval metrics offline, groundedness and answer quality per release, regression gates in CI, and traces that let me reconstruct one production answer from a screenshot. If I cannot replay a bad response, I cannot fix it.
- Architectural trade-offs and engineering decisions. Every answer here is a number: cost per query, tokens per prompt, recall at a latency ceiling. See Data size estimation for how I get those numbers quickly under interview pressure.
The guide says questions will primarily be scenario-based and designed to assess practical problem-solving rather than theoretical memorization. My preparation follows from that: I do not memorize definitions, I rehearse four moves — restate the constraint, name the trade-off, quantify it, then say what would change my answer. Non-AI system design shows up too; Interview data has a rate-limiter race condition, a Redis Lua sliding-window implementation, a Postgres planner ignoring a composite index over 2 million rows, and an active-active two-region collaborative editor with EU data residency and a 120 ms cross-region link. The FDE loop is a full senior engineering loop with an AI layer on top, not an AI trivia loop.
The day-to-day
What I actually do in a week, once hired:- Most hours go into data plumbing that has no glamour: getting the customer’s documents into chunks that survive their formatting, and getting their permissions into my filter.
- A large fraction of the job is explaining and calibrating expectations: showing the customer why recall@5 at 200 ms costs more than recall@20 at 2 s, and which one their use case should pick.
- I write evaluation before scale. A pipeline without a scored eval set cannot be changed safely, and every FDE engagement I have seen accumulates technical debt exactly where evaluation was skipped.
- I am on call for my own demo-to-production delta. The failure modes are known and named above, so the bar is having measured them, not discovering them.
- I ship artefacts other people can run: a repo, a runbook, a dashboard. Leaving a system nobody can operate is a failed engagement regardless of the demo.