Extract structured technical data from PDFs with either GitHub Copilot or local Ollama models. The tool condenses a PDF, identifies and summarizes technologies, then writes validated CSV classifications. It also provides interactive document Q&A, deterministic verification reports, and multi-model benchmarks.
For the usual end-to-end workflow, select a PDF and run:
/extraction
/extraction runs all the dependent stages in sequence:
condense → summarize → classify
The final result is saved to:
02_output/22_tech_classification_csv/<pdf-name>_classification.csv
Run individual stages when you want to inspect or edit the intermediate condesed and summary markdown, before classification:
1. /condense → 01_input/12_condensed_md/<pdf-name>_condensed.md
2. /summarize → 02_output/21_tech_summary_md/<pdf-name>_summary.md
3. /classify → 02_output/22_tech_classification_csv/<pdf-name>_classification.csv
Condense creates a compact cached Markdown version of the PDF. Summarize finds technology names and produces detailed per-technology summaries. Classify converts the summary Markdown into one or more rows per technology, including separate rows for distinct years or time horizons.
Researchers and analysts often need to turn semi-structured technical reports into structured datasets that can be compared across technologies, years, and sources. Manual extraction is labor-intensive, difficult to reproduce, and slow to audit. General-purpose LLM interfaces can assist with individual questions, but they do not preserve a complete, inspectable extraction pipeline.
This software provides a reproducible workflow from PDF input to condensed Markdown, technology-level summaries, structured classification CSVs, and validation reports. Intermediate artifacts remain human-readable and editable, which allows generated content to be reviewed before it becomes structured research data.
# Windows (winget)
winget install Python.Python.3.13# macOS (Homebrew)
brew install pythonpython -m venv .venv
# Windows
.venv\Scripts\python -m pip install -e ".[dev,web]"
# macOS/Linux
.venv/bin/python -m pip install -e ".[dev,web]"Both backends use the same model picker, commands, PDF workflow, output formats, and web interface.
Requires an authenticated GitHub account with an active Copilot subscription.
The Python Copilot SDK includes its compatible CLI runtime. Existing CLI
installations can be selected with COPILOT_CLI_PATH.
Sign in once before running the Copilot backend:
$copilot = Get-ChildItem "$env:LOCALAPPDATA\github-copilot-sdk\cli" -Recurse -Filter copilot.exe |
Sort-Object LastWriteTime -Descending |
Select-Object -First 1
& $copilot.FullName loginThen launch the console or web UI:
python -m copilot download-runtime # optional; first run can download it automatically
techclass # console mode
techclass --web # web UI at http://127.0.0.1:5050If the SDK reports Not authenticated, repeat the login step above or provide
a GitHub token for this PowerShell session:
$env:COPILOT_GITHUB_TOKEN = "<your-token>"When running directly from the repository checkout without activating the editable script, use:
.\.venv\Scripts\python -m techclass --webInstall Ollama, start its server, and pull at least one model:
ollama serve
ollama pull llama3.2
techclass --local # console mode with Ollama
techclass --local --web # web UI with Ollama--ollama is an alias for --local. Ollama mode does not require the Copilot
CLI or a GitHub sign-in. See docs/chat-backends.md for local-backend
settings such as OLLAMA_HOST, OLLAMA_NUM_CTX, OLLAMA_TEMPERATURE, and
OLLAMA_SEED.
The deterministic test suite does not require an LLM backend or network access:
python -m pytestAt startup, select a model and then select or upload a PDF. Use console commands in console mode, or the corresponding buttons in web mode. The web UI supports the main pipeline, its checks, benchmarks, file selection/upload, and document Q&A.
| Command | Description |
|---|---|
/list |
List PDFs in 01_input/11_pdf_to_analyze/ and select one |
/upload <path> |
Copy a PDF into 01_input/11_pdf_to_analyze/ and load it |
/current |
Show the currently selected PDF |
/extraction |
Run the complete pipeline: condense → summarize → classify |
/condense |
Condense the PDF to cached Markdown |
/summarize |
Extract technology summaries to Markdown |
/classify |
Convert the summary Markdown into a structured CSV |
/condense-check |
Check how well numeric data survived PDF condensation |
/batch-analyze <question> |
Ask one question across all PDFs |
/benchmark |
Compare all available models on selected technologies from one PDF |
/commands / /help |
Show all commands |
/exit / /quit |
Exit the console application |
Commands also work without the leading /. An unknown /command is reported as
an error instead of being sent to the model. Other input is treated as a
question for the selected backend. For a selected PDF, condensed text is added
on the first question; follow-up questions reuse the session context.
The tracked Allgoewer_2024 files provide one end-to-end example, from the
source PDF through generated outputs and validation artifacts. The name refers
to the example article by Leo Allgoewer and co-authors; those article authors
are not authors of this software. See examples/README.md
for the file list and source citation.
Run /benchmark, select a PDF, then choose exactly three technologies. The tool
summarizes and classifies those same technologies with every model available
from the selected backend.
Results are written to 02_output/23_validation/benchmark/:
<pdf-name>_<provider>_benchmark_summary_<yyyy-MM-dd>.md— summaries from each model<pdf-name>_<provider>_benchmark_classification_<yyyy-MM-dd>.csv— classified rows with a leadingModelcolumn<pdf-name>_<provider>_benchmark_overview_<yyyy-MM-dd>.csv— timing, word count, row count, and status per model
<provider> is copilot or ollama, matching the active backend.
LLMtool_techClass/
|-- pyproject.toml Python package metadata and dependencies
|-- techclass/ Python package
| |-- chat/ provider-neutral Copilot and Ollama backends
| |-- console/ command-line entry point and command loop
| |-- core/ PDF extraction and pipeline orchestration
| |-- format_output/ records, Markdown, and CSV formatting
| |-- helpers/ parsing, validation, verification, benchmarks
| |-- web/ FastAPI host for the browser interface
| `-- workspace.py shared workspace directory model
|-- tests/ backend-independent Python tests
|-- web/wwwroot/ framework-free browser interface
|-- prompt/ LLM prompt templates
|-- docs/ architecture notes and reviewer guide
|-- 01_input/
| |-- 11_pdf_to_analyze/ input PDFs
| |-- 12_condensed_md/ regenerable condensed Markdown cache
| `-- 13_technology_list_md/ editable frozen technology lists
`-- 02_output/
|-- 21_tech_summary_md/ technology summary Markdown
|-- 22_tech_classification_csv/ classified CSV output
`-- 23_validation/ benchmark and verification reports
- Reviewer guide
- Architecture notes
- Example materials
- Ollama backend notes
- Contribution and support guidelines
- Large PDFs are split into approximately 30 KB chunks before LLM processing.
- Scanned PDFs require OCR preprocessing because extraction uses the PDF text layer.
- Ollama requests have no client-side timeout; stopping the request cancels generation.
- Classification requires a summary Markdown file.
/extractionand/summarizecreate it automatically. - Local Ollama runs default to
temperature=0and a fixed seed for more reproducible results. Copilot does not expose equivalent sampling controls, so values can vary between runs. - Reproducibility means repeatable output under the same conditions; it does not guarantee that extracted values are correct. Review generated summaries and validation reports.
The CSV includes identifiers, descriptions, classification hierarchy, year and location, reference-unit data, maturity and efficiency, input/output carriers and ratios, cost and lifetime values, source publication year, summary, and the model that produced the row when applicable.
If you use this software in research, cite it using CITATION.cff and, once published, the associated JOSS paper. The software is distributed under the MIT License.
- v1.2 — Python port: migrated the project from the legacy C# codebase to a Python implementation and applied minor adjustments.
- v1.1 — Renamed the project to Open-source LLM Tool for Technical Data Extraction and Classification; documented GitHub Copilot and local Ollama backends; added the
/extractionfull workflow and its web UI action. - v1.0 — Data folders restructured into
01_input/(11_pdf_to_analyze,12_condensed_md,13_technology_list_md) and02_output/(21_tech_summary_md,22_tech_classification_csv,23_validation/{benchmark, condensed_md_check, classification_csv_check}); classify verification reports get their own folder; benchmark files named<pdf>_<provider>_benchmark_*_<yyyy-MM-dd>; technology lists renamed<name>_technology_list.md. - v0.6 — Source reorganized by role:
core/,format_output/,chat/,console/, andhelpers/; project renamedTestApp→TechClass. - v0.5 and earlier — Prototype and workflow iterations, including PDF condensation, summary/classification stages, prompt templates, and benchmark support.
