Turning messy audio and photographed documents into structured data without making values up. Speech may be transcribed in the wrong script, OCR misreads digits, a header may be cropped out, and a document that is not a lab report must not produce invented results.
Home / Work / Speech and report extraction that does not invent values
Speech and report extraction that does not invent values
A FastAPI service with speech transcription and structured extraction from photographed lab-report images, built on provider adapters with mock defaults and a parsing layer that does not invent values.
Solo project, extraction. Python, FastAPI, faster-whisper, Tesseract.
- My roleSole contributor in the commit history. All 51 commits, covering routes, schemas, services, adapters, normalizers, tests, containers, documentation and fixtures.
- EvidenceChecked by me Test recordings and synthetic report images are saved in the repository.
Playable system flow
Watch a report line become a row
Pick a line and press Play. The parser runs in your browser, the same port that matches the original Python on every golden line.
The controls need JavaScript. The whole flow is written out below, step by step.
The flow in text
Looks valid, is wrong A misread digit that still parses.
Raw line: A line of report text Recorded project run
- What came in
- The text of one line, as read from a photographed lab report: RP <8.5 mg/dL <1.0 N/A
- What acted
- The OCR reader produced this text; the parser has not touched it.
- What it decided
- Nothing yet.
- What changed
- There is only text. No fields, no numbers.
- What happens next
- The parser splits the line into fields.
A line from the project's test fixtures. The raw OCR output behind it is not committed.
Fields: Split into five fields Runs in browser
- What came in
- The raw line.
- What acted
- The row parser: columns are separated by runs of spaces, or by the collapsed-line pattern.
- What it decided
- Assign the pieces, in order, to test, value, unit, range and flag.
- What changed
- 5 fields found.
- What happens next
- The value field is checked against the known value forms.
- Source
- app/services/report_parser.py (original Python); browser port in src/lib/bench/speech.ts.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Parser rule: Matched: a qualified number Runs in browser
- What came in
- The value field: "<8.5".
- What acted
- The value normaliser tries the known forms in order: a range, a qualified number, a scientific value, a plain number.
- What it decided
- "<8.5" matches a qualified number.
- What changed
- The value becomes a typed number.
- What happens next
- The row is assembled.
- Failure mode
- The value forms are a whitelist. A misread digit that still forms a valid number matches and passes: there is no confidence measure.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Structured row: The structured row Runs in browser
- What came in
- A typed value and four other fields.
- What acted
- The row builder.
- What it decided
- Build the row and keep the raw line beside it.
- What changed
- test "RP", value <8.5, unit mg/dL, range <1.0, flag N/A.
- What happens next
- The result.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Result: Row kept Runs in browser
- What came in
- A valid structured row.
- What acted
- The survival rule: a row survives only if its value parsed.
- What it decided
- Keep it.
- What changed
- The row joins the structured output.
- What happens next
- Compare the output with what the report said.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Meaning: Parse success is not correct meaning Recorded project run
- What came in
- The kept value <8.5 and what the report actually said: <0.5.
- What acted
- A comparison with the original report, from the project README (not something the parser can do).
- What it decided
- The parser accepted a value that is wrong. It has no way to know.
- What changed
- A structured, valid-looking row now carries a wrong value, and nothing flags it.
- What happens next
- End of the trace. The parser is a whitelist, not a verifier of meaning.
- The lesson
- Parse success is not semantic correctness. The output is well-formed and wrong. Catching this needs a confidence measure or a cross-check that does not exist in the project.
Recorded: the project README records this OCR misread (self-reported evaluation); the raw OCR output is not committed.
Caught: row left out A misread that matches no value form.
Raw line: A line of report text Recorded project run
- What came in
- The text of one line, as read from a photographed lab report: Cell Count 1.2 x 1043 10^3/µL 1.0 - 2.0 N/A
- What acted
- The OCR reader produced this text; the parser has not touched it.
- What it decided
- Nothing yet.
- What changed
- There is only text. No fields, no numbers.
- What happens next
- The parser splits the line into fields.
A line from the project's test fixtures. The raw OCR output behind it is not committed.
Fields: Split into five fields Runs in browser
- What came in
- The raw line.
- What acted
- The row parser: columns are separated by runs of spaces, or by the collapsed-line pattern.
- What it decided
- Assign the pieces, in order, to test, value, unit, range and flag.
- What changed
- 5 fields found.
- What happens next
- The value field is checked against the known value forms.
- Source
- app/services/report_parser.py (original Python); browser port in src/lib/bench/speech.ts.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Parser rule: No known value form matches Runs in browser
- What came in
- The value field: "1.2 x 1043".
- What acted
- The value normaliser tries the known forms in order: a range, a qualified number, a scientific value, a plain number.
- What it decided
- "1.2 x 1043" matches none of the four known forms.
- What changed
- There is no value, so the row cannot be built.
- What happens next
- The row is left out.
- Failure mode
- The value forms are a whitelist. A misread digit that still forms a valid number matches and passes: there is no confidence measure.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Structured row: No row is produced Runs in browser
- What came in
- A value that matched no form.
- What acted
- The row builder.
- What it decided
- Omit the row: it is not guessed and not kept.
- What changed
- The report continues without this row.
- What happens next
- The result.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Result: Row left out Runs in browser
- What came in
- No row.
- What acted
- The survival rule: a row survives only if its value parsed.
- What it decided
- Drop it.
- What changed
- The row never reaches the output.
- What happens next
- End of the trace.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
A clean line A plain number.
Raw line: A line of report text Recorded project run
- What came in
- The text of one line, as read from a photographed lab report: Hemoglobin 10.5 g/dL 12.0 - 16.0 L
- What acted
- The OCR reader produced this text; the parser has not touched it.
- What it decided
- Nothing yet.
- What changed
- There is only text. No fields, no numbers.
- What happens next
- The parser splits the line into fields.
A line from the project's test fixtures. The raw OCR output behind it is not committed.
Fields: Split into five fields Runs in browser
- What came in
- The raw line.
- What acted
- The row parser: columns are separated by runs of spaces, or by the collapsed-line pattern.
- What it decided
- Assign the pieces, in order, to test, value, unit, range and flag.
- What changed
- 5 fields found.
- What happens next
- The value field is checked against the known value forms.
- Source
- app/services/report_parser.py (original Python); browser port in src/lib/bench/speech.ts.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Parser rule: Matched: a plain number Runs in browser
- What came in
- The value field: "10.5".
- What acted
- The value normaliser tries the known forms in order: a range, a qualified number, a scientific value, a plain number.
- What it decided
- "10.5" matches a plain number.
- What changed
- The value becomes a typed number.
- What happens next
- The row is assembled.
- Failure mode
- The value forms are a whitelist. A misread digit that still forms a valid number matches and passes: there is no confidence measure.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Structured row: The structured row Runs in browser
- What came in
- A typed value and four other fields.
- What acted
- The row builder.
- What it decided
- Build the row and keep the raw line beside it.
- What changed
- test "Hemoglobin", value 10.5, unit g/dL, range 12.0 - 16.0, flag L.
- What happens next
- The result.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Result: Row kept Runs in browser
- What came in
- A valid structured row.
- What acted
- The survival rule: a row survives only if its value parsed.
- What it decided
- Keep it.
- What changed
- The row joins the structured output.
- What happens next
- End of the trace.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
A qualified value An operator and a number, read correctly.
Raw line: A line of report text Recorded project run
- What came in
- The text of one line, as read from a photographed lab report: CRP <0.5 mg/dL <1.0 N/A
- What acted
- The OCR reader produced this text; the parser has not touched it.
- What it decided
- Nothing yet.
- What changed
- There is only text. No fields, no numbers.
- What happens next
- The parser splits the line into fields.
A line from the project's test fixtures. The raw OCR output behind it is not committed.
Fields: Split into five fields Runs in browser
- What came in
- The raw line.
- What acted
- The row parser: columns are separated by runs of spaces, or by the collapsed-line pattern.
- What it decided
- Assign the pieces, in order, to test, value, unit, range and flag.
- What changed
- 5 fields found.
- What happens next
- The value field is checked against the known value forms.
- Source
- app/services/report_parser.py (original Python); browser port in src/lib/bench/speech.ts.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Parser rule: Matched: a qualified number Runs in browser
- What came in
- The value field: "<0.5".
- What acted
- The value normaliser tries the known forms in order: a range, a qualified number, a scientific value, a plain number.
- What it decided
- "<0.5" matches a qualified number.
- What changed
- The value becomes a typed number.
- What happens next
- The row is assembled.
- Failure mode
- The value forms are a whitelist. A misread digit that still forms a valid number matches and passes: there is no confidence measure.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Structured row: The structured row Runs in browser
- What came in
- A typed value and four other fields.
- What acted
- The row builder.
- What it decided
- Build the row and keep the raw line beside it.
- What changed
- test "CRP", value <0.5, unit mg/dL, range <1.0, flag N/A.
- What happens next
- The result.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Result: Row kept Runs in browser
- What came in
- A valid structured row.
- What acted
- The survival rule: a row survives only if its value parsed.
- What it decided
- Keep it.
- What changed
- The row joins the structured output.
- What happens next
- End of the trace.
Runs in your browser: the parser port, checked against the original Python on 184 golden lines.
Problem
What I built
A service with two workflows. One transcribes Bengali and English speech. The other extracts metadata and result rows from photographed lab-report images, using layout-specific row parsers and value, unit and date normalizers.
Both sit on provider adapters, with a deterministic mock by default and faster-whisper and Tesseract as real providers. The Docker image is mock-only by design.
My contribution
- Role
- Sole contributor in the commit history.
- Personal
- All 51 commits, covering routes, schemas, services, adapters, normalizers, tests, containers, documentation and fixtures.Basis: commit history.
- Team
- None.
- Upstream
- faster-whisper, Tesseract, FastAPI and standard libraries. The test report content is derived from the Synthea synthetic-patient dataset, as documented in the repository.
- Provenance
- Built as a take-home assignment; the assignment brief is not in the repository. All commits fall on a single day (August 2026).
Architecture
- An audio upload goes to the transcription route, the transcription service and a provider adapter (mock or faster-whisper).
- An image upload goes to the document route, the extraction service and an OCR adapter (mock or Tesseract).
- Text lines pass through metadata extraction and layout-specific row parsers, then value, unit and date normalizers.
- A conservative detector returns the type unknown with zero rows for a document that is not a lab report.
OCR
A provider reads a photographed report into text lines. A mock provider is the default and Tesseract is the real one.
- Decision
- Provider adapters with mock defaults.
- Limitation
- OCR degrades on rotated or angled images, and raw outputs are not committed.
Lab report check
A conservative detector decides whether the text looks like a lab report. If not, the type is unknown and zero rows are returned.
- Decision
- Conservative non-lab detection.
- Evidence
- Partially verified A receipt returned the type unknown and 0 rows. Tested on one image; self-reported.
Metadata
Reads patient, date and lab fields where present. With the header cropped out, metadata is null.
- Evidence
- Partially verified With the header cropped out, metadata is null and 4 rows are recovered. Self-reported README tables.
Value parser
Reads the value column as a plain number, a qualified number, a range or a scientific value. Anything that matches none of those forms returns no value.
- Input
- The value text of one row.
- Output
- A value with its kind, operator, numbers and the raw text, or none.
- Why it exists
- A value that does not match a known form is not repaired or guessed.
- Decision
- Reject malformed values, and do not correct OCR using knowledge of the source document.
- Control
- No match returns none, and the row is omitted.
- Failure mode
- A wrong value that is still well formed parses as valid, for example <8.5 read in place of <0.5.
- Limitation
- There is no confidence measure, so the parser cannot detect a misread that is still a valid number.
- Evidence
- Partially verified The README records Tesseract reading CRP <0.5 as RP <8.5 and 1.2 x 10^3 as 1.2 x 1043. Self-reported evaluation; raw OCR outputs are not committed.
- Source
- app/services/value_normalizer.py
Row omission
Drops a row whose value does not parse. A row that does parse keeps its raw line.
- Input
- A parsed value, or none.
- Output
- A result row with its raw line, or no row.
- Why it exists
- A made-up value is worse than a missing row.
- Decision
- Omit rather than guess.
- Control
- A value of none returns no row.
- Failure mode
- An omitted row is silent: the response does not list what was dropped.
- Limitation
- The OCR text of dropped rows is not returned.
- Evidence
- Partially verified Recorded examples: with the header cropped out, 4 rows are recovered; a receipt returns the type unknown and 0 rows. Self-reported README tables.
- Source
- app/services/report_parser.py
Engineering decisions
Recorded in the repository’s decisions document:
- Provider adapters with mock defaults, so the service can be run and tested without models or binaries.
- Structured values that keep operators and both range endpoints.
- Preserve uncertain OCR text, and do not correct it using knowledge of the fixtures.
- Tesseract as the real OCR adapter.
- Conservative non-lab detection.
Evaluation
A manual evaluation I ran of the real providers: eight audio clips (English clean, noisy and low volume; Bengali clean and noisy; code-switched; silence; ambient noise) and eight report images (clean, angled, dark, cropped, rotated, normalization cases and a non-lab receipt). Results are self-reported in README tables; the separate evaluation file in the repository is stale. 75 test functions are defined and none are claimed as passing.
Results
VerifiedTwelve audio files, eight report images, reference transcripts and reports, and a decisions record with five decisions are committed. A separate OCR evaluation file is stale and describes older samples.Read the committed files.
VerifiedThe parsing layer keeps the raw OCR reading, canonicalizes only known unit aliases, and a conservative non-lab detector returns the type unknown with zero rows.Read the source and the decisions record.
Partially verifiedThe README reports rows recovered per report image (9, 8, 7, 4, 7, 5, 1 and 0), and notes that English transcription is good while Bengali and code-switched audio are weak.Read the README tables. Self-reported manual evaluation of the real providers; raw transcription outputs are not committed and the numbers were not reproduced.
Partially verifiedRecorded examples of preserved uncertainty: with the header cropped out, metadata is null and 4 rows are recovered; OCR misreads such as CRP <0.5 read as RP <8.5 are kept as read; a non-lab receipt gives the type unknown and 0 rows.Read the committed inputs and the recorded outcomes. The inputs are committed but the outcome tables are my own; the non-lab case was tested on one image.
Limitations
- The evaluation sets are small and made by me, and the parsing heuristics are tuned to the fixtures.
- Bengali transcription is weak with the chosen model, and OCR degrades on rotated or angled images.
- Transcription orchestration is thin, and there is no lockfile or CI.
- Real-provider results need downloaded models and binaries, and the raw outputs are not committed.
Not claimed Production readiness, scale and performance are not claimed for this project.
Takeaways
Keeping the raw OCR line and omitting rows that cannot be parsed makes errors visible. The cost is partial output, such as 4 rows from a report with its header cropped out.
Repository
- Repository
- No hosted demo.