# Recall & Build interview worksheet

Original practice scenarios. Sources checked 5 October 2026. This file is for your private notes; it does not sync to the app.

For each drill: answer in two minutes, run the linked lab exercise, then revise using evidence. These are practice prompts, not verified employer questions.

## Find where the evidence disappeared

A support answer contradicts a policy you know is in the document collection. What would you inspect before changing the model?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Choose one failing query. Trace its expected passage through ingestion, access filters, candidates, ranking and the final context. Change only the failing stage.

Evidence: A before/after trace with document IDs, the diagnosed stage and a regression case.

Self-review:
- Check the indexed revision and authorized document scope.
- Distinguish a missing candidate from a candidate removed during ranking or context assembly.
- If the passage reaches the model, examine how the answer uses it.

Follow-up: What if the expected document belongs to another customer?

Reference: https://learn.microsoft.com/en-us/azure/search/retrieval-augmented-generation-overview

## Keep the exception with the rule

A refund policy puts its exceptions in a table below the main rule. Your chunks separate them. How will you choose a better split?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Compare two chunk sizes and one structure-aware split on the same small labeled question set. Include a question that needs the table.

Evidence: A comparison of retrieval recall, answer support and context tokens, plus one failed example.

Self-review:
- Preserve useful document structure and source identifiers.
- Treat size and overlap as experiment settings, not universal constants.
- Keep evaluation questions fixed while comparing alternatives.

Follow-up: Could increasing overlap improve recall while making the answer more expensive?

Reference: https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-chunk-documents

## Find both product codes and paraphrases

Exact product-code queries work well, but descriptions in everyday language miss the right page. Should you add vectors, fusion or reranking?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Use the retrieval lab to compare lexical, dense and fused rankings. Add a reranker only as a measured extension.

Evidence: Recall and ranking results split by exact identifiers and paraphrases, with retrieval latency.

Self-review:
- Explain the different signals supplied by text and vector search.
- Keep fusion and reranking distinct.
- A reranker cannot rescue a relevant document absent from its candidates.

Follow-up: When would the extra ranking step be too costly?

Reference: https://learn.microsoft.com/en-us/azure/search/hybrid-search-overview

## Justify the autonomy

A ticket process has fixed validation and approval steps, but some requests need unpredictable information gathering. Which parts should an agent control?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Draw a fixed workflow and an agent variant for the same support task. Run representative cases using your selected provider’s existing agent lab.

Evidence: A decision note comparing success, number of calls, permissions and failure handling.

Self-review:
- Keep known business rules in application code.
- Name the uncertainty that justifies dynamic tool selection.
- Define explicit stopping and escalation conditions.

Follow-up: What would make you replace the agent with a workflow?

Reference: https://www.anthropic.com/engineering/building-effective-agents

## Stop a loop without repeating an action

A ticket tool times out after writing successfully. The agent repeatedly calls it because it never saw confirmation. How should the application recover?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: In a local fixture, lose a successful tool response. Add bounded attempts and reconcile the stored result using a stable operation ID before retrying.

Evidence: A failure trace, a stop reason and an assertion that only one ticket was created.

Self-review:
- Separate a failed operation from an unknown result.
- Bound elapsed time, turns and spending independently.
- Check the system of record before repeating a consequential action.

Follow-up: What should the learner see if reconciliation is also unavailable?

Reference: https://www.anthropic.com/engineering/building-effective-agents
Reference: https://developers.openai.com/api/docs/guides/agent-builder-safety

## Decide whether a change can ship

A new prompt improves average answers but fails more often when no supporting document exists. Would you release it?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Add answerable, unsupported and adversarial slices to the evaluation lab. Compare the baseline and candidate with predetermined release criteria.

Evidence: A slice-level report including failures, latency, cost and a written release decision.

Self-review:
- Separate development examples from the held-out test set.
- Combine executable checks with calibrated human or model judgments.
- Inspect repeated trials and important failure slices rather than relying on one average.

Follow-up: How would you detect that your judge rewards longer answers?

Reference: https://developers.openai.com/api/docs/guides/evaluation-best-practices

## Check the action behind the answer

The agent says it resolved a ticket and its final message looks correct. Its trace shows a forbidden tool attempt. Is the run successful?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Replay allowed and disallowed tool paths in the agent lab. Check tool arguments, approvals and stored outcomes separately from the final text.

Evidence: A trace checklist and tests for an allowed action, a blocked action and an unsupported success claim.

Self-review:
- Define permitted steps as well as the required final state.
- Score tool choice, arguments and approval compliance.
- Verify completion in the target system; a fluent claim is insufficient.

Follow-up: Can a blocked unsafe attempt still count as a quality failure?

Reference: https://developers.openai.com/api/docs/guides/trace-grading
Reference: https://developers.openai.com/api/docs/guides/agent-builder-safety

## Validate meaning after validating shape

A model returns a perfectly valid ticket object containing a nonexistent account and an unsupported priority. What can your API safely accept?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Extend the ticket lab with schema-valid but business-invalid fixtures. Exercise refusal, missing output and timeout paths.

Evidence: Boundary tests showing which inputs are rejected and that no unauthorized action runs.

Self-review:
- Use schema validation for shape and separate checks for business constraints.
- Keep authorization outside the model’s decision.
- Handle refusals and incomplete output explicitly.

Follow-up: Which fields should trusted server context supply instead of the model?

Reference: https://developers.openai.com/api/docs/guides/structured-outputs
Reference: https://developers.openai.com/api/docs/guides/agent-builder-safety

## Optimize the slow stage

Your support agent meets its quality target, but tail latency and cost per resolved ticket are too high. Which change do you test first?

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Measure retrieval, queue, model and tool time. Compare one routing or call-reduction change on the same evaluation set. Use the cache lab for a separate cache experiment.

Evidence: A baseline/candidate report for task success, p95 latency and cost per successful task, including retries.

Self-review:
- Find the bottleneck before choosing an optimization.
- Parallelize only independent work and distinguish streaming responsiveness from completion time.
- Check quality after reducing context, calls or model size.

Follow-up: Which requests can tolerate deferred batch processing, and which cannot?

Reference: https://developers.openai.com/api/docs/guides/latency-optimization
Reference: https://developers.openai.com/api/docs/guides/evaluation-best-practices

## Explain a failure you can prove

A stakeholder asks what you personally built, what failed and how you know the fix helped. Prepare an answer backed by your project artifacts.

- My initial answer:
- My assumptions:
- What I built or changed:
- Failing case and diagnosis:
- Commit / test output / trace:
- Measurements before and after:
- What remains uncertain:
- My revised answer:

Exercise: Choose one real lab failure. Record the symptom, baseline, investigation, smallest fix and rerun. Label local simulations honestly.

Evidence: A two-minute explanation linking your commit, failing case and before/after measurements.

Self-review:
- Separate your contribution from team work.
- State whether the evidence is local, simulated, deployed or from real users.
- Name one unresolved limitation and the next useful experiment.

Follow-up: What would you do if a launch decision had to be made with incomplete evidence?

Reference: https://developers.openai.com/api/docs/guides/evaluation-best-practices
