Imtiaz Mashrafee

Home / Work / Two-stage email drafting pipeline

Two-stage email drafting pipeline

A Python project that compares a single-prompt email generator with a two-stage pipeline that validates inputs before generating, scored on ten recorded scenarios.

Solo project, LLM pipeline. Python, Streamlit, LLM provider APIs, Pydantic.

  • My roleSole contributor in the commit history. The design of both approaches, the input checker, the metric framework, the ten-scenario set and the Streamlit demo.
  • EvidenceChecked by me Ten test scenarios are saved in the repository, with the recorded outputs and the scores a judge model gave them.

Playable system flow

Watch a request travel the pipeline

Pick a recorded scenario and press Play. Validation runs in your browser. The model steps are recorded outcomes of the project run.

The controls need JavaScript. The whole flow is written out below, step by step.

The flow in text

Vague request: asks first The checker holds the request back and asks.

  1. Request: A request arrives Recorded project run

    What came in
    A request: "I need an email about my job assessment.", with 1 fact and a professional tone.
    What acted
    The request enters the pipeline.
    What it decided
    Nothing yet. No model has been called.
    What changed
    The request is now the pipeline input.
    What happens next
    Validation checks the request in code.

    Scenario 3 of the project's ten committed scenarios.

  2. Validation: Validation, in code Runs in browser

    What came in
    The request fields: intent, key facts, tone.
    What acted
    Request validation (a typed schema plus a non-empty check), before any model is called.
    What it decided
    The request has the basics, so it may go on.
    What changed
    Nothing. The request passes unchanged.
    What happens next
    The input checker (a model) judges whether there is enough to write from.

    The ported validator runs on this request in your browser. Its messages are tested against the original Python.

  3. Input checker: The input checker judges Recorded project run

    What came in
    The validated request.
    What acted
    The input checker, a model call that decides whether there is enough usable information.
    What it decided
    Not enough to write from, so clarification is needed.
    What changed
    The status becomes "clarification".
    What happens next
    The pipeline returns a question and writes no email.

    A model call decided this. It is not run in this page. The checker's own output was not saved, so this decision is read from the recorded outcome (it asked, or it wrote). Why it decided so is not known.

  4. Clarify or generate: Clarify: no email is written Recorded project run

    What came in
    A "clarification" status.
    What acted
    The clarify or generate branch.
    What it decided
    Return a question and stop.
    What changed
    No email is produced. The clarification text was not kept in the record.
    What happens next
    The evaluation compares this outcome with what was expected.

    Recorded outcome of the project run. The branch runs inside the pipeline, not in this page.

  5. Evaluation: Evaluation: as expected Recorded project run

    What came in
    The outcome (asked first) and the scenario's expected behavior (ask first).
    What acted
    The offline evaluation runner. It is not part of the live pipeline: it scores both strategies on the ten scenarios afterwards.
    What it decided
    The outcome matches what was expected.
    What changed
    Recorded judge scores out of 5: a single prompt 2.33, the checked pipeline 5.
    What happens next
    End of the run. Reset to replay, or inspect any step.
    Limitation
    Ten scenarios and one run. The checker is a model and in one recorded scenario it did not ask when it should have.

    Recorded. One run, ten scenarios, scored by a model whose identity was not recorded.

Recorded miss The checker should have asked, and did not.

  1. Request: A request arrives Recorded project run

    What came in
    A request: "Request a deadline extension for my final project.", with 2 facts and a urgent but respectful tone.
    What acted
    The request enters the pipeline.
    What it decided
    Nothing yet. No model has been called.
    What changed
    The request is now the pipeline input.
    What happens next
    Validation checks the request in code.

    Scenario 4 of the project's ten committed scenarios.

  2. Validation: Validation, in code Runs in browser

    What came in
    The request fields: intent, key facts, tone.
    What acted
    Request validation (a typed schema plus a non-empty check), before any model is called.
    What it decided
    The request has the basics, so it may go on.
    What changed
    Nothing. The request passes unchanged.
    What happens next
    The input checker (a model) judges whether there is enough to write from.

    The ported validator runs on this request in your browser. Its messages are tested against the original Python.

  3. Input checker: The input checker judges Recorded project run

    What came in
    The validated request.
    What acted
    The input checker, a model call that decides whether there is enough usable information.
    What it decided
    Enough to write from, so generation can go ahead.
    What changed
    The request goes through as clean input.
    What happens next
    The pipeline generates the email.

    A model call decided this. It is not run in this page. The checker's own output was not saved, so this decision is read from the recorded outcome (it asked, or it wrote). Why it decided so is not known.

  4. Clarify or generate: Generate: the email is written Recorded project run

    What came in
    Clean input and a "go ahead" status.
    What acted
    The clarify or generate branch.
    What it decided
    Continue to the generator.
    What changed
    The generator is called.
    What happens next
    The email is generated.

    Recorded outcome of the project run. The branch runs inside the pipeline, not in this page.

  5. Generation: Generation Recorded project run

    What came in
    Clean input from the checker.
    What acted
    The generator, a model call.
    What it decided
    None. It writes.
    What changed
    An email is produced. The recorded run did not keep its text, so none is shown.
    What happens next
    The evaluation scores it against what was expected.

    Recorded: the run produced an email. The text was not saved.

  6. Evaluation: Evaluation: a recorded miss Recorded project run

    What came in
    The outcome (wrote the email) and the scenario's expected behavior (ask first).
    What acted
    The offline evaluation runner. It is not part of the live pipeline: it scores both strategies on the ten scenarios afterwards.
    What it decided
    The outcome does not match: this request should have triggered clarification.
    What changed
    Recorded judge scores out of 5: a single prompt 2, the checked pipeline 2.33.
    What happens next
    End of the run. Reset to replay, or inspect any step.
    Limitation
    Ten scenarios and one run. The checker is a model and in one recorded scenario it did not ask when it should have.

    Recorded. One run, ten scenarios, scored by a model whose identity was not recorded.

Clear request: writes Enough information, so the pipeline writes.

  1. Request: A request arrives Recorded project run

    What came in
    A request: "Follow up after today's project sync to recap what we agreed and confirm the next steps.", with 3 facts and a friendly tone.
    What acted
    The request enters the pipeline.
    What it decided
    Nothing yet. No model has been called.
    What changed
    The request is now the pipeline input.
    What happens next
    Validation checks the request in code.

    Scenario 1 of the project's ten committed scenarios.

  2. Validation: Validation, in code Runs in browser

    What came in
    The request fields: intent, key facts, tone.
    What acted
    Request validation (a typed schema plus a non-empty check), before any model is called.
    What it decided
    The request has the basics, so it may go on.
    What changed
    Nothing. The request passes unchanged.
    What happens next
    The input checker (a model) judges whether there is enough to write from.

    The ported validator runs on this request in your browser. Its messages are tested against the original Python.

  3. Input checker: The input checker judges Recorded project run

    What came in
    The validated request.
    What acted
    The input checker, a model call that decides whether there is enough usable information.
    What it decided
    Enough to write from, so generation can go ahead.
    What changed
    The request goes through as clean input.
    What happens next
    The pipeline generates the email.

    A model call decided this. It is not run in this page. The checker's own output was not saved, so this decision is read from the recorded outcome (it asked, or it wrote). Why it decided so is not known.

  4. Clarify or generate: Generate: the email is written Recorded project run

    What came in
    Clean input and a "go ahead" status.
    What acted
    The clarify or generate branch.
    What it decided
    Continue to the generator.
    What changed
    The generator is called.
    What happens next
    The email is generated.

    Recorded outcome of the project run. The branch runs inside the pipeline, not in this page.

  5. Generation: Generation Recorded project run

    What came in
    Clean input from the checker.
    What acted
    The generator, a model call.
    What it decided
    None. It writes.
    What changed
    An email is produced. The recorded run did not keep its text, so none is shown.
    What happens next
    The evaluation scores it against what was expected.

    Recorded: the run produced an email. The text was not saved.

  6. Evaluation: Evaluation: as expected Recorded project run

    What came in
    The outcome (wrote the email) and the scenario's expected behavior (write the email).
    What acted
    The offline evaluation runner. It is not part of the live pipeline: it scores both strategies on the ten scenarios afterwards.
    What it decided
    The outcome matches what was expected.
    What changed
    No per-scenario judge score is quoted here.
    What happens next
    End of the run. Reset to replay, or inspect any step.
    Limitation
    Ten scenarios and one run. The checker is a model and in one recorded scenario it did not ask when it should have.

    Recorded. One run, ten scenarios, scored by a model whose identity was not recorded.

Validation stops it Scenario 3 with every fact removed. Runs here.

  1. Request: The same request, with no facts Runs in browser

    What came in
    "I need an email about my job assessment.", with the key facts removed.
    What acted
    You (this is an edit of scenario 3, not a recorded scenario).
    What it decided
    Nothing yet.
    What changed
    The request has no facts.
    What happens next
    Validation checks it in code.

    Built in your browser from scenario 3 by removing its facts.

  2. Validation: Validation rejects it Runs in browser

    What came in
    A request with an empty fact list.
    What acted
    Request validation, in code.
    What it decided
    Rejected: List should have at least 1 item after validation, not 0.
    What changed
    The pipeline stops here. No model is called.
    What happens next
    Nothing: the input checker is never reached.

    The ported validator runs on this request in your browser. Its messages are tested against the original Python.

Problem

A single prompt that drafts an email goes ahead even when the request is unclear, contradictory, padded with irrelevant facts or contains instruction-like text, and then writes a confident email.

The project asks whether a validation stage before generation changes that behavior, and how to measure the difference.

What I built

Two approaches to the same task. Model A is a single prompt that writes the email. Model B is a two-stage pipeline: an input checker that canonicalizes the intent, filters facts, records excluded bullets and detects contradictions and ambiguity, then either returns a clarification question or hands a clean input to the generator.

Around them: a provider wrapper, schemas and prompts, an evaluation runner with three judge metrics and structural checks, and a Streamlit app that drives the pipeline.

My contribution

Role
Sole contributor in the commit history.
Personal
The design of both approaches, the input checker, the metric framework, the ten-scenario set and the Streamlit demo.Basis: commit history.
Team
None.
Upstream
LLM provider APIs (a Gemini default and an NVIDIA-compatible endpoint) and standard libraries.
Provenance
The repository has a single bulk initial commit (June 2026), so no development history is visible.

Architecture

Email validation gateA clear request passes validation, generates an email, which the evaluation runner can later judge. An ambiguous request takes the ask first branch and generates nothing.inputvalidatewrites the email
The check passes and the generator writes the email.The check returns a clarification question; nothing is generated.Input, the validation gate with an ask-first branch, and the generator; the judge belongs to the evaluation runner.Modelled from the code
  1. Inputs (intent, key facts, tone) are validated as a request.
  2. The input checker produces a status of OK or CLARIFY, with the canonical intent, kept facts and excluded bullets.
  3. On CLARIFY, the pipeline returns a question. On OK, the generator writes the email.
  4. The evaluation runner sends scenarios through both strategies, scores each result with three judge metrics and structural checks, and saves JSON.
  1. Request validation

    Checks the request in code before any model is called: the intent is not empty, there is at least one non-empty fact, and the tone is one of the allowed values.

    Input
    Intent, key facts and tone.
    Output
    A valid request, or an error that stops the run.
    Why it exists
    Cheap, deterministic failures should never reach a model.
    Decision
    Validate in code first.
    Control
    A validation error is raised before the input checker runs.
    Failure mode
    Code can only check shape. It cannot tell whether the request is clear enough.
    Limitation
    Whether a request is clear is decided by a model, not by this stage.
    Evidence
    Verified The messages and rules shown here match the original Python on a grid of test inputs. Checked with the pydantic version recorded in the golden file.
    Source
    src/schemas.py, model_b_approach/pipeline.py
  2. Input checker

    A model call that decides whether the request has enough usable information. It returns a status, a canonical intent, the usable facts and, when needed, a clarifying question.

    Input
    A valid request.
    Output
    Ready with a canonical intent and usable facts, or clarify with a question.
    Why it exists
    Ask first instead of guessing.
    Decision
    Make clarification a normal outcome of the pipeline.
    Control
    A clarify status returns the question and no email is generated.
    Failure mode
    It is a model, so it can miss: in one recorded scenario a missing extension length did not trigger a question.
    Limitation
    Not run in this page. Outcomes are recorded, and any edit retires them.
    Evidence
    Partially verified Ten scenarios, one run: the gated pipeline asked first in 2 of the 3 scenarios that expected a question. Judge and generator models were not recorded; the result was not reproduced.
    Source
    model_b_approach/input_checker.py
  3. Clarify or generate

    On a clarify status the pipeline returns the question and no email. Otherwise the generator writes the email.

    Decision
    Make clarification a normal outcome of the pipeline.
    Source
    model_b_approach/pipeline.py
  4. Evaluation runner

    A separate runner sends scenarios through both strategies and scores the results with three judge metrics and structural checks. It is not part of the pipeline.

    Evidence
    Partially verified The committed summary averages are 4.7 for the gated pipeline and 4.2 for the baseline, on a scale of 1 to 5. Ten scenarios, one run; the judge and generator models are not recorded.
    Source
    evaluate.py

Engineering decisions

  • Separate validation from generation, so ambiguity can end the run with a question instead of a guess.
  • Record what was excluded and what must be clarified, instead of silently dropping input.
  • Three metrics with scores from 1 to 5 and written justifications: intent and fact fidelity, ambiguity and clarification handling, and quality and delivery readiness.
  • Judge at temperature 0, with JSON replies, retries and score clamping.
  • Scenarios designed around specific behaviors (clean request, wordy intent, vague purpose, missing critical fact, irrelevant filler, injection attempt, contradiction, uncertainty to preserve, apology, concise update).

Evaluation

Ten hand-written scenarios, two strategies, three LLM-judge metrics scored from 1 to 5 with justifications, plus deterministic structural checks on the email. One run is saved to JSON. The scenarios are synthetic.

Results

VerifiedThe ten-scenario set with expected behaviors, and a results file of 20 scored records (two strategies across ten scenarios), are committed.Read the committed files and compared them with the source.

Partially verifiedThe committed summary averages are 4.7 for the gated pipeline and 4.2 for the baseline, on a scale of 1 to 5.Read the committed results. Ten scenarios, one run; the judge and generator models are not recorded and the result was not reproduced.

VerifiedIn the code, a validation stage returns either OK or CLARIFY before any email is generated.Read the source.

Partially verifiedRecorded scenario outcomes: for a vague assessment email the baseline wrote an email (overall 2.33) and the gated pipeline asked for clarification (5.0); for a deadline extension with no length, both strategies wrote an email and the gate did not trigger (baseline 2.0, gated 2.33); for contradictory timing both asked for clarification (baseline 5.0, gated 4.67).Read the committed per-scenario records. One run, one judge; stage internals were not saved, so why the gate passed or triggered cannot be shown without a re-run.

The gated pipeline helped in some scenarios, matched the baseline in others and missed in one. The averages are not presented as proof that it is better in general.

Limitations

  • Ten scenarios and one run; the judge and generator models were not recorded and the LLM judge is not calibrated.
  • Stage internals were not saved, so why the gate triggered or passed in a given scenario cannot be shown without a re-run.
  • There are no automated tests, and re-running the evaluation needs LLM API keys.
  • The metric definition document linked from the README is missing.

Not claimed Production readiness, scale and performance are not claimed for this project.

Takeaways

Making clarification a possible outcome gave the evaluation something to score besides the email itself.

The recorded miss, where the gate did not trigger, is the reason the averages are shown with their qualifier and not as a headline.

Repository