Paper Claims Pipeline¶
The paper claims pipeline extracts verifiable technical claims from PDF papers. It is a reusable single-document
operation under osa_tool.operations.analysis.paper_claims and is not registered in the legacy scheduler or agent graph.
Use osa-tool --paper-analysis when those claims must be verified against a repository. Add
--include-repository-quality when the same run also needs the formal repository-quality score.
The current flow is:
Status¶
This standalone extraction stage focuses on conversion, section parsing, claim extraction, and
local evaluation utilities. The paper_analysis CLI supplies the standard repository-specific integration for the
typed result objects.
Runtime behavior¶
- PDFs are split with
pypdfinto physical chunks before Marker conversion. - The default chunk size is five pages and can be changed per run.
- Temporary chunk PDFs are deleted after conversion.
- Marker Markdown is cached under the system temporary directory.
- Section parsing and LLM claim extraction are intentionally rerun every time.
- LLM responses are validated with Pydantic.
- Invalid claim candidates are repaired through the repair prompt; after the final retry, bad claim candidates are dropped so one bad claim does not fail the whole document.
original_textis checked against the source section using exact matching first, then conservative RapidFuzz matching.- Claims are checked for plausible language script against their source evidence and section context.
Public Python API¶
from pathlib import Path
from osa_tool.operations.analysis.paper_claims import PaperClaimPipeline, PipelineOptions
pipeline = PaperClaimPipeline(model_handler)
result = await pipeline.arun(Path("paper.pdf"), PipelineOptions(pages_per_chunk=5))
The synchronous wrapper is available for scripts:
The main public objects are:
| Object | Purpose |
|---|---|
PdfChunker |
Validates and splits PDFs into temporary physical chunks. |
MarkerDocumentConverter |
Converts PDF chunks through Marker and caches successful Markdown output. |
MarkdownSectionParser |
Parses merged Markdown into ordered PaperSection objects. |
ClaimExtractor |
Runs section selection, per-section claim extraction, and deduplication. |
PaperClaimPipeline |
Composes the single-document pipeline. |
PipelineResult |
Holds converted Markdown, sections, and typed extraction results. |
clear_marker_cache |
Deletes Marker cache entries. |
Exported artifacts¶
PaperClaimPipeline.export(...) writes:
| File | Description |
|---|---|
document.md |
Merged Marker Markdown. |
sections.json |
Parsed sections with heading metadata. |
claims.json |
Typed extraction schema when legacy=False. |
claims_legacy.json |
MVP-compatible claim JSON when legacy=True. |
report.json |
Canonical stage report with paper source, configured model, actual successful models, and typed extraction result. |
Legacy JSON excludes debug-only step3_selection by default:
Use include_debug=True when you need deduplication debug data:
payload = result.to_legacy_dict(include_debug=True)
PaperClaimPipeline.export(result, "out/paper", legacy=True, include_debug=True)
The debug payload is stored under:
Batch CLI¶
The batch utility processes one or more PDF files or directories through the single-document pipeline:
Useful options:
| Option | Default | Description |
|---|---|---|
--chunk-pages |
5 |
Number of PDF pages per physical chunk. |
--max-retries |
5 |
LLM response validation and repair attempts. |
--dedup-batch-size |
50 |
Maximum extracted claims sent in one deduplication request. |
--model |
openai/gpt-5.4-mini |
Model name passed through the normal OSA validation model settings. |
--include-debug |
false |
Include debug-only data such as debug.step3_selection in claims_legacy.json. |
--force-marker-refresh |
false |
Ignore cached Marker Markdown and reconvert PDFs. LLM extraction is still rerun. |
--marker-process-isolation / --no-marker-process-isolation |
true |
Run each Marker chunk in a separate Python process to release CUDA memory between chunks. |
--marker-low-vram / --no-marker-low-vram |
true |
Use conservative Marker batch sizes for low-VRAM GPUs. |
--marker-log-cuda-memory / --no-marker-log-cuda-memory |
true |
Log CUDA memory before and after each Marker chunk when CUDA is available. |
Example for a low-VRAM GPU:
python -m osa_tool.tools.paper_claims.batch ./papers \
--chunk-pages 5 \
--marker-low-vram \
--marker-process-isolation
Example with debug selection output:
Marker cache¶
Only successful Marker Markdown is cached. The cache key includes the PDF content hash, chunk size, Marker version, and relevant conversion options. Incomplete cache entries are ignored and reconverted.
Use --force-marker-refresh from the CLI or MarkerOptions(force_refresh=True) from Python to bypass the cache for one
run.
To delete cache entries programmatically:
Evaluation utilities¶
Install the paper-claims extra before running the conversion or evaluation utilities:
Python compatibility: core OSA supports Python 3.11 and later, but the PDF-to-claims conversion workflow requires Python 3.11--3.14. Marker and its dependency stack are not currently available for Python 3.15+. On those interpreters, the default converter fails early with an actionable compatibility error instead of attempting conversion with an incomplete extra.
Run semantic matching:
python -m osa_tool.tools.paper_claims.evaluate \
--llm paper_claim_results/paper/claims_legacy.json \
--human annotations.json \
--output evaluation.json
Aggregate evaluation outputs:
Dependencies¶
The paper-claims extra provides pypdf, markdown-it-py, rapidfuzz, marker-pdf, numpy, scipy, and
sentence-transformers. Pandas remains part of OSA's core dependencies and is used by the aggregate utility.
Marker, Markdown parsing, and RapidFuzz validation are loaded lazily when their corresponding pipeline stage runs. OSA does not enable Marker's LLM processors.
Module layout¶
The reusable operation lives in:
Important modules:
| Module | Responsibility |
|---|---|
models.py |
Pydantic data contracts and legacy serialization. |
pdf_splitter.py |
PDF validation and physical chunk creation. |
marker_converter.py |
Marker conversion, cache handling, and low-VRAM/process-isolated execution. |
section_parser.py |
Markdown-to-section parsing. |
claim_schemas.py |
Private Pydantic schemas for LLM response validation. |
claim_validation.py |
Source-text matching, script guard, and claim candidate partitioning. |
claim_input_planner.py |
Token budgets plus hierarchy-aware selection and sentence-aware claim inputs. |
claim_deduplicator.py |
Token-bounded, fail-safe LLM claim deduplication. |
claim_extractor.py |
LLM request/repair loop and three-step extraction orchestration. |
pipeline.py |
Single-document pipeline composition and artifact export. |
Supporting command-line tools live in: