Citations Manager for the Sly.so Blog
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Balazs Horvath 1c81dd51e6 Implement ignore_domains, which the README documented but nothing read
The config parser stored `ignore_domains` and no consumer consulted it, so
the documented escape hatch for bot-filtering sites silently did nothing:
adding cnbc.com changed nothing about the health report.

check_health now excludes ignored domains from the error count, and records
them under `ignored` plus `ignored_details`. They are deliberately not
dropped: a site that filters our requests must not be able to look healthy
by accident, so the count stays visible in both the report and the exit
status derivation. Host matching is anchored on a dot so a suffix lookalike
like notcnbc.com is not silently swallowed by cnbc.com.

Repository.health and the check/health subcommands thread the config value
through. HealthReport gains the `ignored` field, and `healthy` stops being
computed over ignored entries, which otherwise stayed in the error list and
kept the report unhealthy regardless of the new count.

Two related silent-failure fixes while in these paths:

- The two blind `except Exception` handlers in the validator turned any
  unexpected error, including a bug in the validator itself, into
  "unreachable", so a code fault was indistinguishable from a dead link.
  The broad catch is kept deliberately (a validator must not abort a run
  over one URL) but now prefixes "unexpected:" so the two are tellable
  apart. The existing test asserting only `error is not None` still holds.
- Repository.load discarded parse failures with a bare `pass`, so a
  malformed .bib file made citations vanish with no trace. Failures are
  now collected in `load_errors` and printed by `citar health`.

Import ordering and blank-line fixes across the package are incidental to
getting the tree lint-clean, which the edit gate required first.
2026-10-04 12:49:51 +02:00
citar Implement ignore_domains, which the README documented but nothing read 2026-10-04 12:49:51 +02:00
tests Implement ignore_domains, which the README documented but nothing read 2026-10-04 12:49:51 +02:00
.gitignore citar v0.1.0: reusable citation manager with BibTeX parse, URL validation, HTML footnote generation, and Pelican plugin 2026-07-29 09:59:11 +02:00
pyproject.toml citar v0.1.0: reusable citation manager with BibTeX parse, URL validation, HTML footnote generation, and Pelican plugin 2026-07-29 09:59:11 +02:00
pytest.ini feat(citar): skip @String/@Comment/@Preamble in bibtex parser + add comprehensive pytests 2026-07-29 11:11:44 +02:00
README.md citar v0.1.0: reusable citation manager with BibTeX parse, URL validation, HTML footnote generation, and Pelican plugin 2026-07-29 09:59:11 +02:00

citar — Citations Manager for the Sly.so Blog

Reusable citation/references manager that parses BibTeX files from ~/research/, validates citation links, and generates Pelican-ready HTML footnotes for blog posts.

Installation

cd ~/src/citar
pip install -e .

Requires requests (already in the environment). No other external deps.

Usage

CLI

# Build a merged citation database from all ~/research/*.bib files
citar build

# Save to a JSON file for inspection
citar build -o ~/.citar/citations.json

# Check all citation URLs are reachable
citar check

# Generate Markdown citations for a post (inline + references section)
citar cite -k arxiv:2607.10183v2,spiritbuun/buun-llama-cpp

# Generate HTML inline citations (for posts with raw HTML)
citar cite -k arxiv:2607.10183v2 -f html

# Search citations by keyword
citar search DSPark
citar search turbo kv

# Health check with full report
citar health

Python API

from citar import build_repository, body_citation, references_section_html

repo = build_repository(["~/research"])

# Inline citations in order
body = repo.format_body_html(["arxiv:2607.10183v2", "spiritbuun/buun-llama-cpp"])
# → '<a href="#ref-1">[1]</a> <a href="#ref-2">[2]</a>'

# References section at the bottom
refs = repo.format_references_html(["arxiv:2607.10183v2", "spiritbuun/buun-llama-cpp"])

# Health check
report = repo.health()
print(report.healthy)

Configuration

Create ~/.citar.toml:

[paths]
bib_dirs = "~/research"
citation_dir = "~/.citar"
validation_timeout = 10

Integration with Pelican Blog

See blog_plugin.py in the repo root — a Pelican custom content processor that auto-generates footnote HTML in ## References sections.

Alopex Integration

The validator borrows concepts from Alopex's verification pipeline:

  • Multi-step validation (HEAD → GET fallback)
  • Graceful degradation on timeouts and connection errors
  • Clear error categorization (timeout, connection error, too many redirects, HTTP status)

Format

Inline Citations (body)

<a href="#ref-16">[16]</a>

Reference Items (bottom)

<p><a id="ref-16"></a><a href="https://arxiv.org/abs/2607.10183">[16]</a>
Paper: "Automated Tensor Scheduling..." (arXiv:2607.10183)</p>

Data Models

  • Citation — parsed BibTeX entry with title, author, year, URL, DOI
  • ValidationResult — single URL check result (reachable, status code, error)
  • HealthReport — aggregate health check for the full citation set