🔬 Impresso Datalab Notebooks
-
Updated
Aug 20, 2026 - Jupyter Notebook
🔬 Impresso Datalab Notebooks
A Toponym Resolution Pipeline for Digitised Historical Newspapers
This repository provides underlying code and materials for the paper 'Station to Station: Linking and Enriching Historical British Railway Data'.
An experimental infrastructure for computational critique of historical writing using RAG architectures, argumentation analysis, rhetorical mapping, and large language models.
Official repository of WikiTextGraph
Teaching materials for the workshop "Network Analysis in the Humanities" organised by DigiTS (Center for Digital Text Scholarship) at the University of Tartu.
Training classifier models to predict genres and subgenres on album cover data.
RISHI-Q — computational comparative history of physics. Flagship: Vaiśeṣika ākāśa–śabda sound-medium ontology vs Greek/Chinese/Buddhist controls & Maxwell EM (9/9 · 6/6 · 0/5 · R2 unique).
Training classifier models on acoustic metadata to predict genres and subgenres.
Replication package for PHTS Theory v3.0 — a two-tier structural framework for the Voynich Manuscript. Submitted to Cryptologia, 2026. Contains Python scripts, data files, and pre-computed results.
Word–color association from large-scale online image data (CIELch). Code for the PLOS ONE manuscript.
A machine-checked formalization of the Fuxi 64-hexagram system. 40 claims verified; 7 corrected, including a clustering coefficient that is 0 not 5/12, and a divination kernel whose stationary distribution is not uniform.
Impresso Python Library to interact with the Impresso Public API
5D semantic-dynamics framework separating Korean prose poetry from flash fiction. Code for the Scientific Reports manuscript.
Chunk-level FAISS + Solar-10.7B RAG pipeline over 2,900+ Korean flash fiction texts.
Computational vocabulary analysis across Buddhist Vajrayana, Shakta Tantra and Baul Bengali texts · TF-IDF char n-grams · cosine similarity · Tara to Kali lexical migration · Ramprasad Sen OCR · Charyapada · 75 texts across 8 tradition layers
A source-grounded, falsifiable atlas of humanity's shared story-patterns: 338 public-domain sacred texts, 106 traditions, 52k quote-anchored motif records, preregistered cross-cultural universality tests.
Information dynamics in Korean flash fiction — surprisal, coherence, and semantic-shift trajectories. Code for the Physica A manuscript.
Readability (LIX) and lexical diversity (TTR/MATTR) with the parameters other implementations hardcode
To associate your repository with the computational-humanities topic, visit your repo's landing page and select "manage topics."