Skip to content

About

A tool to track the bibliography of my field

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

94 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

BiblioAssistant is an automated pipeline designed for keeping track of new scientific literature. It filters high volumes of daily publications and generates high-value summaries for researchers, with a focus on hydrology, climate, and meteorology. The selection is specifically tailored to the research interests of the Hydrology and Climate Change research group at the Ebro Observatory.

Features

  • Waterfall Filter Architecture:
    • Ingestion: Monitors RSS feeds and OpenAlex for author publications, citations, and journal updates.
    • Relevance Filtering: High-throughput local processing using Ollama (recommended: llama3.1:8b) to maintain privacy, eliminate API costs, and evaluate hundreds of abstracts daily without quota limits.
    • Synthesis: Deep synthesis of identified relevant papers into structured "Extended Cards" using Google Gemini API (gemini-2.5-flash). Because only 2–6 papers are typically synthesized per day, this fits comfortably within Gemini's free tier (up to 20 requests/day at 0€ cost) while providing 1M token context window and near-instant processing (5–10s per paper) without local CPU timeouts.
  • Full-Text Extraction: Automated PDF download and text extraction (with HTML fallback).
  • Static Site Generation: Beautiful, bookish-style website for browsing summaries.
  • MathJax Support: High-quality rendering of LaTeX equations.
  • RSS Feed: Dedicated feed for the generated summaries.
  • Automated Deployment: Easy rsync-based deployment to remote servers.

Prerequisites

  • Python 3.12+
  • uv (Python package and project manager)
  • Ollama (for local LLM processing)
  • Optional / Recommended: Google Gemini API Key from Google AI Studio for synthesis

Setup

  1. Clone the repository:

    git clone <repository-url>
    cd biblioassistant
  2. Install dependencies: Using uv:

    uv sync
  3. Configure Ollama: Pull the recommended local model for relevance filtering:

    ollama pull llama3.1:8b
  4. Environment Variables: Copy .env.example to .env and adjust your configuration:

    cp .env.example .env

    Key configuration settings:

    # Relevance Filtering Engine (Local Ollama is recommended)
    RELEVANCE_ENGINE=ollama
    OLLAMA_FILTER_MODEL=llama3.1:8b
    OLLAMA_HOST=http://localhost:11434
    
    # Synthesis Engine (Gemini API is recommended)
    SYNTHESIS_ENGINE=gemini-api
    GEMINI_API_KEY=your_gemini_api_key_here
    GEMINI_MODEL=gemini-2.5-flash
    
    # Ollama Fallback / Local Synthesis Model
    OLLAMA_MODEL=gemma4:31b
    
    # Budget Control
    MAX_MONTHLY_COST=10.0                 # Maximum monthly spend in Euro
    
    # Deployment configuration
    REMOTE_HOST=your.server.com
    REMOTE_USER=your_username
    REMOTE_PATH=/var/www/biblio/
    
    # OpenAlex Polite Pool
    OPENALEX_EMAIL=your-email@example.com

Usage

Run the full pipeline (Discovery -> Filter -> Synthesize -> Generate -> Deploy):

uv run python -m src.main --deploy

Command Line Arguments

  • --deploy: Sync the generated site to the remote server.
  • --force-all: Ignore the "seen" database and re-process all entries.
  • --generate-only: Skip fetching and processing; only rebuild the static site.
  • --add-doi <DOI>: Manually add a specific paper by DOI.
  • --backfill <days>: Set the start date for discovery to N days ago.
  • --to-date <YYYY-MM-DD>: Set the end date for discovery (useful for backfilling).
  • --backfill-mode: Set the "added date" to the paper's publication date. This prevents historical papers from appearing in the RSS feed or the "Recent" section.

Utility Scripts

  • Add Authors by DOI: Expand the monitored authors list by fetching all authors from a specific paper.
    uv run scripts/add_authors_by_doi.py <DOI>
    This script resolves the DOI via OpenAlex, extracts all author IDs, and adds them to the monitored_authors database table.

Historical Backfilling

BiblioAssistant includes a mechanism to progressively populate its database with historical papers without overwhelming the RSS feed or the main page.

The backfill.py script:

  1. Retrieves a "cursor" from the database (starting 7 days ago if first run).
  2. Processes a 7-day window of historical papers.
  3. Sets their entry date to their publication date (--backfill-mode).
  4. Moves the cursor back by 7 days for the next run.
  5. Stops automatically when it reaches January 1, 2000.

This script is automatically called by run_daily.sh after the main pipeline run. To run it manually:

uv run python backfill.py --deploy

Scheduling

To run the pipeline automatically every day, you have two options on Linux:

Option 1: Standard Cron (User-level)

Ideal for servers that are always on.

  1. Open your crontab: crontab -e
  2. Add the following line:
    0 6 * * * cd /path/to/biblioassistant && /usr/local/bin/uv run python -m src.main --deploy >> /path/to/biblioassistant/data/cron.log 2>&1

Option 2: Cron Daily (System-level / Recommended for Laptops)

This uses anacron to ensure the job runs even if the computer was off at the scheduled time.

  1. Create a launcher script: sudo nano /etc/cron.daily/biblioassistant
  2. Paste the following:
    #!/bin/sh
    # Launcher for BiblioAssistant
    su your_username -c "/path/to/biblioassistant/run_daily.sh"
  3. Make it executable: sudo chmod +x /etc/cron.daily/biblioassistant

The system uses the run_daily.sh script provided in the repository to manage the execution environment.

Project Structure

  • src/: Core Python modules.
  • templates/: Jinja2 templates for the static site.
  • data/: Local storage for the database, PDFs, and Markdown summaries (ignored by git).
  • public/: The generated static website (ignored by git).

Development Philosophy

Separation of Engine and Content: This repository is designed to contain only the "engine" of the project: the source code, configuration structures, and templates. All "content" (processed data, PDFs, markdown summaries, databases, and news items) must reside in the data/ directory, which is excluded from version control.

This ensures that the repository remains lightweight and focuses on the software logic, while the data is managed as a separate, local-first artifact. Files such as data/news.json or data/db.sqlite3 should never be force-added to the repository.

Author

Developed by Pere Quintana Seguí.

This project was partially funded by Fundació Observatori de l'Ebre.

This project was developed with the assistance of AI tools, specifically Gemini CLI.

License

GPLv3

About

A tool to track the bibliography of my field

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages