Answering questions over research papers, including tables, with answers that are checked rather than trusted, and with repeated work made cheaper through caching.
Home / Work / Retrieval over research papers with verify and repair
Retrieval over research papers with verify and repair
A Python system for question answering over parsed research PDFs, with hybrid retrieval, a verify-and-repair step, layered caches and a harness that records each evaluated query's path.
Solo project, retrieval. Python, FastAPI, DuckDB, sentence-transformers.
- My roleSole contributor in the commit history. All six commits in the repository, which hold the ingestion, retrieval, generation, verification, caching, API and evaluation code.
- EvidenceChecked by me Saved evaluation runs record which path each question took through the pipeline.
Playable system flow
Watch a question pass the quality gate
Pick a recorded question and press Play. The pass or repair rule runs in your browser. Only the route, the repair flag and the final score were recorded; the pipeline stages are described from the code, and the drafts came from a mock generator.
The controls need JavaScript. The whole flow is written out below, step by step.
The flow in text
Repaired, still below (q3) Repair ran; recorded final score 0.67.
Question: A question arrives Recorded project run
- What came in
- "Compare CNNs and Transformers"
- What acted
- The API hands the question to the retrieval pipeline.
- What it decided
- Nothing yet.
- What changed
- The question becomes the pipeline input.
- What happens next
- A query plan is chosen from its intent.
Query q3 of the committed evaluation set (20 queries, eight recorded runs).
Plan: A plan is chosen Modelled from the code
- What came in
- The question.
- What acted
- The query planner.
- What it decided
- A plan is chosen from the question's intent, then executed step by step.
- What changed
- The pipeline now has an ordered list of steps to run.
- What happens next
- Hybrid retrieval finds candidate passages.
Described from the implementation. Which plan ran for this query was not recorded.
Retrieve: Hybrid retrieval Modelled from the code
- What came in
- The question and the plan.
- What acted
- Two indexes behind an ensemble retriever: BM25 and an in-memory embedding store.
- What it decided
- Weight the two equally the first time (BM25 0.5, vectors 0.5) and take the top 12.
- What changed
- Up to 12 candidate passages (the actual count for this query was not recorded).
- What happens next
- The candidates are reranked.
Described from the implementation and its default configuration.
Rerank: Rerank Modelled from the code
- What came in
- The candidate passages.
- What acted
- A deterministic re-sort by the retrieval score. It is not a second relevance signal or a learned reranker.
- What it decided
- Keep the best 8.
- What changed
- Up to 8 passages, best first.
- What happens next
- Irrelevant sentences are pruned and the context is built.
Described from the implementation.
Prune and contextualize: Prune and build the context Modelled from the code
- What came in
- The reranked passages.
- What acted
- A sentence-level prune, then context construction.
- What it decided
- Drop sentences that share fewer than 1 word with the question (a lexical overlap test, not a relevance judgment); cap the context.
- What changed
- A context of at most 4000 characters.
- What happens next
- A draft answer is generated.
Described from the implementation and its default configuration.
Draft: A draft is written Recorded project run
- What came in
- The question and the context.
- What acted
- The generator, at temperature 0.2.
- What it decided
- None. It writes.
- What changed
- A draft answer exists. The committed runs used a mock generator and do not include answer text, so none is shown.
- What happens next
- The draft is scored.
Recorded: the committed runs use a mock generator, so scores describe the harness, not answer quality.
Quality score: What the quality score is made of Modelled from the code
- What came in
- The draft and the question.
- What acted
- The verifier, a rule-based scorer.
- What it decided
- The draft must score at least 0.72 to be accepted.
- What changed
- Nothing yet: this step shows the formula, not a result.
- What happens next
- The score is computed.
- What it measures
- A rule-based answer-quality score from 0 to 1 (higher is better): length, how many of the question's words the draft uses, sentence structure, completeness and repetition.
- What it does not measure
- It is not proof that the answer is grounded in the retrieved sources. The verifier is handed the context but does not use it, and it cannot tell whether a claim is true.
- Score weights
- Length 20%, Relevance 30%, Coherence 25%, Completeness 15%, Repetition 10%. Per-query values for each part were not recorded.
- Where it lives
- rag_papers/generation/verifier.py, ResponseVerifier.verify_response; threshold in Stage4Config.accept_threshold (default 0.72).
From the verifier's code. The five parts of the score were not recorded per query.
Quality score: The first draft is scored Recorded project run
- What came in
- The first draft.
- What acted
- The verifier.
- What it decided
- The first draft fell below 0.72: the run recorded that repair ran, which only happens below the threshold.
- What changed
- The first draft is not good enough.
- What happens next
- The pass-or-repair rule routes the run to repair.
Recorded: repair ran. The first draft's actual score was not recorded, so it is not shown.
Pass or repair: Below the line: repair Recorded project run
- What came in
- A first draft below 0.72.
- What acted
- The acceptance rule.
- What it decided
- Below the required score, and one repair attempt is allowed, so repair runs.
- What changed
- Route: repair. This is the only repair the run gets (maximum 1).
- What happens next
- The repair step tightens the settings and runs again.
Recorded: this query took the repair route. The comparison itself is shown at the final step.
Repair: One bounded repair attempt Modelled from the code
- What came in
- The question and the failed draft.
- What acted
- The repair step: retrieve again with tighter settings, prune harder, regenerate with a constrained template.
- What it decided
- Change five settings, once. There is no retry loop.
- What changed
- BM25 weight 0.5 to 0.7; vector weight 0.5 to 0.3; prune overlap 1 to 2; template compare to constrained; temperature 0.2 to 0.1.
- What happens next
- The new draft is scored again.
- Failure mode
- After the single attempt the answer can still be below the threshold and is then returned as not accepted.
The settings are the router configuration. What the new draft contained was not recorded.
Quality score: The second score: 0.67 Recorded project run
- What came in
- The repaired draft.
- What acted
- The verifier again.
- What it decided
- The repaired draft scored 0.67.
- What changed
- A recorded score exists for the final draft.
- What happens next
- The final comparison.
The recorded final score after repair, for this query.
Final state: Returned: marked not accepted Runs in browser
- What came in
- A final score of 0.67 against 0.72.
- What acted
- The acceptance rule, run on the final score.
- What it decided
- 0.67 is below 0.72. Repair is not repeated, so the answer is returned marked not accepted.
- What changed
- The pipeline returns the answer with an accepted flag of false.
- What happens next
- End of the run. Repair controlled the workflow; it did not guarantee quality.
- The lesson
- Bounded repair controls workflow behavior. The verifier is a heuristic, not a grounding check, so a pass or a fail says little about whether the answer is right.
The final comparison runs in your browser on the recorded final score.
Passes first time Recorded score 0.82.
Question: A question arrives Recorded project run
- What came in
- "What is transfer learning?"
- What acted
- The API hands the question to the retrieval pipeline.
- What it decided
- Nothing yet.
- What changed
- The question becomes the pipeline input.
- What happens next
- A query plan is chosen from its intent.
Query q1 of the committed evaluation set (20 queries, eight recorded runs).
Plan: A plan is chosen Modelled from the code
- What came in
- The question.
- What acted
- The query planner.
- What it decided
- A plan is chosen from the question's intent, then executed step by step.
- What changed
- The pipeline now has an ordered list of steps to run.
- What happens next
- Hybrid retrieval finds candidate passages.
Described from the implementation. Which plan ran for this query was not recorded.
Retrieve: Hybrid retrieval Modelled from the code
- What came in
- The question and the plan.
- What acted
- Two indexes behind an ensemble retriever: BM25 and an in-memory embedding store.
- What it decided
- Weight the two equally the first time (BM25 0.5, vectors 0.5) and take the top 12.
- What changed
- Up to 12 candidate passages (the actual count for this query was not recorded).
- What happens next
- The candidates are reranked.
Described from the implementation and its default configuration.
Rerank: Rerank Modelled from the code
- What came in
- The candidate passages.
- What acted
- A deterministic re-sort by the retrieval score. It is not a second relevance signal or a learned reranker.
- What it decided
- Keep the best 8.
- What changed
- Up to 8 passages, best first.
- What happens next
- Irrelevant sentences are pruned and the context is built.
Described from the implementation.
Prune and contextualize: Prune and build the context Modelled from the code
- What came in
- The reranked passages.
- What acted
- A sentence-level prune, then context construction.
- What it decided
- Drop sentences that share fewer than 1 word with the question (a lexical overlap test, not a relevance judgment); cap the context.
- What changed
- A context of at most 4000 characters.
- What happens next
- A draft answer is generated.
Described from the implementation and its default configuration.
Draft: A draft is written Recorded project run
- What came in
- The question and the context.
- What acted
- The generator, at temperature 0.2.
- What it decided
- None. It writes.
- What changed
- A draft answer exists. The committed runs used a mock generator and do not include answer text, so none is shown.
- What happens next
- The draft is scored.
Recorded: the committed runs use a mock generator, so scores describe the harness, not answer quality.
Quality score: What the quality score is made of Modelled from the code
- What came in
- The draft and the question.
- What acted
- The verifier, a rule-based scorer.
- What it decided
- The draft must score at least 0.72 to be accepted.
- What changed
- Nothing yet: this step shows the formula, not a result.
- What happens next
- The score is computed.
- What it measures
- A rule-based answer-quality score from 0 to 1 (higher is better): length, how many of the question's words the draft uses, sentence structure, completeness and repetition.
- What it does not measure
- It is not proof that the answer is grounded in the retrieved sources. The verifier is handed the context but does not use it, and it cannot tell whether a claim is true.
- Score weights
- Length 20%, Relevance 30%, Coherence 25%, Completeness 15%, Repetition 10%. Per-query values for each part were not recorded.
- Where it lives
- rag_papers/generation/verifier.py, ResponseVerifier.verify_response; threshold in Stage4Config.accept_threshold (default 0.72).
From the verifier's code. The five parts of the score were not recorded per query.
Quality score: The score: 0.82 Recorded project run
- What came in
- The draft.
- What acted
- The verifier.
- What it decided
- The draft scored 0.82.
- What changed
- The pipeline has a number to compare with the requirement.
- What happens next
- The pass-or-repair rule compares it with the required score.
The recorded final score for this query in the committed runs.
Pass or repair: At or above the line: accepted Runs in browser
- What came in
- A score of 0.82 and a required score of 0.72.
- What acted
- The acceptance rule (the ported router logic).
- What it decided
- 0.82 is above 0.72, so the answer is accepted and repair is skipped.
- What changed
- Route: accept. No repair runs.
- What happens next
- The answer is returned.
This comparison runs in your browser, in the same function the route map uses.
Final state: Returned: accepted Recorded project run
- What came in
- An accepted draft.
- What acted
- The API.
- What it decided
- None.
- What changed
- The answer goes back to the caller as accepted.
- What happens next
- End of the run.
Recorded: this query was accepted without repair.
Exactly at the line Recorded score 0.72.
Question: A question arrives Recorded project run
- What came in
- "Summarize the key concepts in neural network training"
- What acted
- The API hands the question to the retrieval pipeline.
- What it decided
- Nothing yet.
- What changed
- The question becomes the pipeline input.
- What happens next
- A query plan is chosen from its intent.
Query q5 of the committed evaluation set (20 queries, eight recorded runs).
Plan: A plan is chosen Modelled from the code
- What came in
- The question.
- What acted
- The query planner.
- What it decided
- A plan is chosen from the question's intent, then executed step by step.
- What changed
- The pipeline now has an ordered list of steps to run.
- What happens next
- Hybrid retrieval finds candidate passages.
Described from the implementation. Which plan ran for this query was not recorded.
Retrieve: Hybrid retrieval Modelled from the code
- What came in
- The question and the plan.
- What acted
- Two indexes behind an ensemble retriever: BM25 and an in-memory embedding store.
- What it decided
- Weight the two equally the first time (BM25 0.5, vectors 0.5) and take the top 12.
- What changed
- Up to 12 candidate passages (the actual count for this query was not recorded).
- What happens next
- The candidates are reranked.
Described from the implementation and its default configuration.
Rerank: Rerank Modelled from the code
- What came in
- The candidate passages.
- What acted
- A deterministic re-sort by the retrieval score. It is not a second relevance signal or a learned reranker.
- What it decided
- Keep the best 8.
- What changed
- Up to 8 passages, best first.
- What happens next
- Irrelevant sentences are pruned and the context is built.
Described from the implementation.
Prune and contextualize: Prune and build the context Modelled from the code
- What came in
- The reranked passages.
- What acted
- A sentence-level prune, then context construction.
- What it decided
- Drop sentences that share fewer than 1 word with the question (a lexical overlap test, not a relevance judgment); cap the context.
- What changed
- A context of at most 4000 characters.
- What happens next
- A draft answer is generated.
Described from the implementation and its default configuration.
Draft: A draft is written Recorded project run
- What came in
- The question and the context.
- What acted
- The generator, at temperature 0.2.
- What it decided
- None. It writes.
- What changed
- A draft answer exists. The committed runs used a mock generator and do not include answer text, so none is shown.
- What happens next
- The draft is scored.
Recorded: the committed runs use a mock generator, so scores describe the harness, not answer quality.
Quality score: What the quality score is made of Modelled from the code
- What came in
- The draft and the question.
- What acted
- The verifier, a rule-based scorer.
- What it decided
- The draft must score at least 0.72 to be accepted.
- What changed
- Nothing yet: this step shows the formula, not a result.
- What happens next
- The score is computed.
- What it measures
- A rule-based answer-quality score from 0 to 1 (higher is better): length, how many of the question's words the draft uses, sentence structure, completeness and repetition.
- What it does not measure
- It is not proof that the answer is grounded in the retrieved sources. The verifier is handed the context but does not use it, and it cannot tell whether a claim is true.
- Score weights
- Length 20%, Relevance 30%, Coherence 25%, Completeness 15%, Repetition 10%. Per-query values for each part were not recorded.
- Where it lives
- rag_papers/generation/verifier.py, ResponseVerifier.verify_response; threshold in Stage4Config.accept_threshold (default 0.72).
From the verifier's code. The five parts of the score were not recorded per query.
Quality score: The score: 0.72 Recorded project run
- What came in
- The draft.
- What acted
- The verifier.
- What it decided
- The draft scored 0.72.
- What changed
- The pipeline has a number to compare with the requirement.
- What happens next
- The pass-or-repair rule compares it with the required score.
The recorded final score for this query in the committed runs.
Pass or repair: Exactly on the line: accepted Runs in browser
- What came in
- A score of 0.72 and a required score of 0.72.
- What acted
- The acceptance rule (the ported router logic).
- What it decided
- 0.72 is equal to 0.72, so the answer is accepted and repair is skipped. The comparison is inclusive: a score equal to the threshold passes.
- What changed
- Route: accept. No repair runs.
- What happens next
- The answer is returned.
This comparison runs in your browser, in the same function the route map uses.
Final state: Returned: accepted Recorded project run
- What came in
- An accepted draft.
- What acted
- The API.
- What it decided
- None.
- What changed
- The answer goes back to the caller as accepted.
- What happens next
- End of the run.
Recorded: this query was accepted without repair.
Problem
What I built
An ingestion path that parses PDFs, chunks them and normalizes tables into DuckDB and Parquet. Two indexes (BM25 and an in-memory embedding store using cosine similarity) behind an ensemble retriever. A query plan chosen from the question’s intent, then executed as retrieve, rerank, prune, contextualize, generate, verify and repair.
Around the pipeline: SQLite stores and three cache layers, asynchronous ingestion jobs, chat history and memory modules, a FastAPI service, command-line tools, a Streamlit app and an evaluation runner.
My contribution
- Role
- Sole contributor in the commit history.
- Personal
- All six commits in the repository, which hold the ingestion, retrieval, generation, verification, caching, API and evaluation code.Basis: commit history.
- Team
- None.
- Upstream
- Third-party libraries and models, including a PDF-parsing wrapper, sentence-transformers embeddings, DuckDB, FastAPI and Streamlit.
- Provenance
- Most of the code arrived in a few very large commits (five in October 2025, one in December 2025), so the development history is not visible. Scaffold placeholder metadata remains in the repository.
Architecture
- PDFs are parsed, chunked and indexed (lexical and vector).
- A query is turned into an intent plan, and the plan is run: retrieve, rerank, prune at sentence level, contextualize, generate.
- The verifier scores the draft. A score at or above the threshold skips repair. A lower score sends the run through one repair attempt, and the answer is then returned either way, marked not accepted if the score is still below the threshold.
- Embedding, retrieval and answer results are cached in three layers.
- The evaluation runner records each query’s path, score and latency to CSV, Parquet and Markdown.
Query plan
A plan is chosen from the question's intent and then executed.
- Decision
- A plan keyed on query intent instead of one fixed chain of steps.
Hybrid retrieval
Two indexes, BM25 and an in-memory embedding store using cosine similarity, behind an ensemble retriever.
- Decision
- Hybrid lexical and vector retrieval.
Sentence-level prune
Irrelevant sentences are pruned before the context is built.
- Decision
- Sentence-level pruning before contextualization, so the generator sees less irrelevant text.
Acceptance rule
Compares the verify score of a draft answer with the acceptance threshold. At or above the threshold the answer is accepted and repair is skipped.
- Input
- A verify score from 0 to 1 and the threshold (0.72 by default).
- Output
- Accepted, or one repair attempt.
- Why it exists
- Acceptance is an explicit rule with a number, not an assumption.
- Decision
- Make acceptance a threshold that a request can override.
- Control
- A score at or above the threshold is accepted; below it, one repair runs.
- Failure mode
- The verify score is only as meaningful as the verifier and the generator; the committed runs use a mock generator.
- Limitation
- No claim is made about answer quality or real model behavior.
- Evidence
- Verified Committed runs record the final score, whether repair ran and whether the answer was accepted. Mock generator; scores describe harness behavior, not answer quality.
- Source
- rag_papers/retrieval/router_dag.py, run_plan and exec_repair
Repair step
Makes one more pass with tighter settings: more weight on BM25 (0.5 to 0.7) and less on vectors (0.5 to 0.3), a stricter prune, a constrained template and a lower temperature (0.2 to 0.1), then verifies again.
- Input
- A draft below the threshold.
- Output
- A new draft and a new verify score.
- Why it exists
- A bounded second attempt has a fixed cost, and acceptance is decided by the same rule.
- Decision
- One repair only.
- Control
- The maximum number of repairs is 1.
- Failure mode
- After the single attempt the answer can still be below the threshold and is then returned as not accepted.
- Limitation
- Whether a repair improves an answer is not claimed.
- Evidence
- Partially verified In the committed runs that produced scores, every query that used repair ended below 0.72 and was not accepted. Mock generator; scores describe harness behavior.
- Source
- rag_papers/retrieval/router_dag.py, exec_repair
Engineering decisions
- A plan keyed on query intent instead of one fixed chain of steps.
- Sentence-level pruning before contextualization, so the generator sees less irrelevant text.
- A verify-then-repair step, so a low-scoring draft gets one more pass, not a retry loop.
- Hybrid lexical and vector retrieval.
- A generator abstraction with a mock default, so tests and demos run offline.
- An evaluation harness that records the path, not only the final answer.
Evaluation
Eight committed runs of a 20-query set, produced by the harness with the mock generator. Metrics include Accept@1, average verify score, repair rate and latency. Test functions are defined (308 unit, 43 integration) and none are claimed as passing.
Results
VerifiedEight committed evaluation runs of a 20-query set record each query's pipeline path, verify score, acceptance flag and whether repair was used.Read the committed run files.
Partially verifiedFigures from the committed runs: Accept@1 50%, average verify score 0.709, repair rate 50% and mean latency of about 9.5 ms with a mock generator.Read the committed summaries. The generator is a mock, so these figures describe harness behavior, not answer quality.
VerifiedThe code defaults to a mock generator, and the committed runs were produced with it, which is why per-query answers are generic.Read the source and the run outputs.
Not claimedThe README states a 400 times answer-cache speedup. In the repository the figure is repeated in README prose, status documents and example-script text, with no retained timing benchmark, so it is not claimed here.Searched the repository for timing artifacts and found none.
Limitations
- The default generator is a mock, so committed evaluation figures describe pipeline plumbing, not answer quality.
- No real-model or real-corpus evaluation is committed, so the full path from PDF to answer with a real model is not demonstrated.
- Chat paths contain TODOs and placeholder sources, two integration test files are empty and the README status is stale.
- Test functions are defined (308 unit and 43 integration) but passing is not claimed, and there is no CI.
Not claimed Production readiness, scale and performance are not claimed for this project.
Takeaways
A metric is only as informative as the component behind it. The harness reported Accept@1 of 50% while the answers were generic, because the generator was a mock, so the figures are labeled as plumbing evidence.
A claim in a README needs an artifact. The speedup figure has none, so it stays out of this page.
Repository
- Repository
- No hosted demo.