Skip to content

Paper Claims Pipeline

The paper claims pipeline extracts verifiable technical claims from PDF papers. It is a reusable single-document operation under osa_tool.operations.analysis.paper_claims and is not registered in the legacy scheduler or agent graph. Use osa-tool --paper-analysis when those claims must be verified against a repository. Add --include-repository-quality when the same run also needs the formal repository-quality score.

The current flow is:

PDF → physical PDF chunks → Marker Markdown → structured sections → extracted claims

Status

This standalone extraction stage focuses on conversion, section parsing, claim extraction, and local evaluation utilities. The paper_analysis CLI supplies the standard repository-specific integration for the typed result objects.

Runtime behavior

  • PDFs are split with pypdf into physical chunks before Marker conversion.
  • The default chunk size is five pages and can be changed per run.
  • Temporary chunk PDFs are deleted after conversion.
  • Marker Markdown is cached under the system temporary directory.
  • Section parsing and LLM claim extraction are intentionally rerun every time.
  • LLM responses are validated with Pydantic.
  • Invalid claim candidates are repaired through the repair prompt; after the final retry, bad claim candidates are dropped so one bad claim does not fail the whole document.
  • original_text is checked against the source section using exact matching first, then conservative RapidFuzz matching.
  • Claims are checked for plausible language script against their source evidence and section context.

Public Python API

from pathlib import Path

from osa_tool.operations.analysis.paper_claims import PaperClaimPipeline, PipelineOptions

pipeline = PaperClaimPipeline(model_handler)
result = await pipeline.arun(Path("paper.pdf"), PipelineOptions(pages_per_chunk=5))

The synchronous wrapper is available for scripts:

result = pipeline.run(Path("paper.pdf"))

The main public objects are:

Object Purpose
PdfChunker Validates and splits PDFs into temporary physical chunks.
MarkerDocumentConverter Converts PDF chunks through Marker and caches successful Markdown output.
MarkdownSectionParser Parses merged Markdown into ordered PaperSection objects.
ClaimExtractor Runs section selection, per-section claim extraction, and deduplication.
PaperClaimPipeline Composes the single-document pipeline.
PipelineResult Holds converted Markdown, sections, and typed extraction results.
clear_marker_cache Deletes Marker cache entries.

Exported artifacts

PaperClaimPipeline.export(...) writes:

File Description
document.md Merged Marker Markdown.
sections.json Parsed sections with heading metadata.
claims.json Typed extraction schema when legacy=False.
claims_legacy.json MVP-compatible claim JSON when legacy=True.
report.json Canonical stage report with paper source, configured model, actual successful models, and typed extraction result.

Legacy JSON excludes debug-only step3_selection by default:

payload = result.to_legacy_dict()

Use include_debug=True when you need deduplication debug data:

payload = result.to_legacy_dict(include_debug=True)
PaperClaimPipeline.export(result, "out/paper", legacy=True, include_debug=True)

The debug payload is stored under:

{
  "debug": {
    "step3_selection": []
  }
}

Batch CLI

The batch utility processes one or more PDF files or directories through the single-document pipeline:

python -m osa_tool.tools.paper_claims.batch ./paper.pdf --output-dir paper_claim_results

Useful options:

Option Default Description
--chunk-pages 5 Number of PDF pages per physical chunk.
--max-retries 5 LLM response validation and repair attempts.
--dedup-batch-size 50 Maximum extracted claims sent in one deduplication request.
--model openai/gpt-5.4-mini Model name passed through the normal OSA validation model settings.
--include-debug false Include debug-only data such as debug.step3_selection in claims_legacy.json.
--force-marker-refresh false Ignore cached Marker Markdown and reconvert PDFs. LLM extraction is still rerun.
--marker-process-isolation / --no-marker-process-isolation true Run each Marker chunk in a separate Python process to release CUDA memory between chunks.
--marker-low-vram / --no-marker-low-vram true Use conservative Marker batch sizes for low-VRAM GPUs.
--marker-log-cuda-memory / --no-marker-log-cuda-memory true Log CUDA memory before and after each Marker chunk when CUDA is available.

Example for a low-VRAM GPU:

python -m osa_tool.tools.paper_claims.batch ./papers \
  --chunk-pages 5 \
  --marker-low-vram \
  --marker-process-isolation

Example with debug selection output:

python -m osa_tool.tools.paper_claims.batch ./paper.pdf --include-debug

Marker cache

Only successful Marker Markdown is cached. The cache key includes the PDF content hash, chunk size, Marker version, and relevant conversion options. Incomplete cache entries are ignored and reconverted.

Use --force-marker-refresh from the CLI or MarkerOptions(force_refresh=True) from Python to bypass the cache for one run.

To delete cache entries programmatically:

from osa_tool.operations.analysis.paper_claims import clear_marker_cache

clear_marker_cache()

Evaluation utilities

Install the paper-claims extra before running the conversion or evaluation utilities:

pip install "osa_tool[paper-claims]"

Python compatibility: core OSA supports Python 3.11 and later, but the PDF-to-claims conversion workflow requires Python 3.11--3.14. Marker and its dependency stack are not currently available for Python 3.15+. On those interpreters, the default converter fails early with an actionable compatibility error instead of attempting conversion with an incomplete extra.

Run semantic matching:

python -m osa_tool.tools.paper_claims.evaluate \
  --llm paper_claim_results/paper/claims_legacy.json \
  --human annotations.json \
  --output evaluation.json

Aggregate evaluation outputs:

python -m osa_tool.tools.paper_claims.aggregate ./evaluations --output aggregate.csv

Dependencies

The paper-claims extra provides pypdf, markdown-it-py, rapidfuzz, marker-pdf, numpy, scipy, and sentence-transformers. Pandas remains part of OSA's core dependencies and is used by the aggregate utility.

Marker, Markdown parsing, and RapidFuzz validation are loaded lazily when their corresponding pipeline stage runs. OSA does not enable Marker's LLM processors.

Module layout

The reusable operation lives in:

osa_tool/operations/analysis/paper_claims/

Important modules:

Module Responsibility
models.py Pydantic data contracts and legacy serialization.
pdf_splitter.py PDF validation and physical chunk creation.
marker_converter.py Marker conversion, cache handling, and low-VRAM/process-isolated execution.
section_parser.py Markdown-to-section parsing.
claim_schemas.py Private Pydantic schemas for LLM response validation.
claim_validation.py Source-text matching, script guard, and claim candidate partitioning.
claim_input_planner.py Token budgets plus hierarchy-aware selection and sentence-aware claim inputs.
claim_deduplicator.py Token-bounded, fail-safe LLM claim deduplication.
claim_extractor.py LLM request/repair loop and three-step extraction orchestration.
pipeline.py Single-document pipeline composition and artifact export.

Supporting command-line tools live in:

osa_tool/tools/paper_claims/