Industry Served: EdTech
An End-to-End Content Intelligence Platform for Publishers, Broadcasters, and Research Institutions
Druid Learning is a Dublin-founded technology company that transforms decades of archives, content libraries, and information repositories into AI-ready data. Publishers, broadcasters, and research institutions sit on enormous proprietary content estates: backlist titles, scanned books, audio-visual files, legacy flash material, PDFs, and metadata scattered across a dozen systems. Very few have the infrastructure to safely optimise and capitalise on it.
We developed the entire platform: a data intelligence layer that ingests content from any source, uses machine learning and natural language processing to enrich and standardise its metadata, and activates the resulting proprietary dataset inside the customer's existing analytics, editorial, and AI applications. Crucially, the source content is never altered and data ownership never leaves the client.
The platform now serves organisations including NPR, CJ Fallon, the Dutch Cancer Institute, and the University of Bristol, and is backed by Techstars, Enterprise Ireland, the European Union's Horizon programme, and LvlUp Ventures. Druid Learning was named winner of Frankfurter Buchmesse's 2026 Wildcard.
Content sat on FTP servers, in legacy databases, and behind APIs, as OCR, XML, INDD, MP4, PDF, EPUB, and legacy flash files.
Itemising blocks of digital content and tagging component assets was manual, slow, and expensive to manage at scale.
Libraries acquired through acquisition or built on author-submitted records had no consistent, usable schema.
A single PDF could hold images, text, tables, and links with no means of separating them into individually addressable, reusable assets.
We built a scalable, modular Content Intelligence Platform enabling:
The platform is delivered as connected products sitting on one pipeline, so clients can start by indexing and categorising an archive and expand into analytics, editorial suggestions, and metadata standardisation without changing their IT infrastructure or existing content.
The highest level of administration, allowing the Druid Learning team to:
A workspace for each publishing house, university, or institution, enabling client authorities to:
The core working environment, built around an Index menu covering Add File, Indexing History, Saved Projects, and Create New Project. Content creators can:
Admin defines courses, levels, sessions, and onboards each learning centre and DOS.
Machine learning labels every file by content topic, cross-labels for relevancy, and semantically links related content across the archive.
Creators index a chosen range of pages rather than reprocessing entire documents, with full indexing history retained.
Assets organised into projects and folders by topic area, level or grade, and subject, then exported at asset, folder, or project level.
Sample datasets validated with subject-matter experts before scaling, built into the process rather than bolted on.
Source content is never altered and clients retain full ownership and control of their dataset.
500k+
Assets Processed
100k+
Data Points Generated
140+
Languages Supported
100%
Of Your Data, Your Contro
Public Broadcasting
Audio, video, and editorial archives consolidated and transformed into a labelled, AI ready dataset. 500k+ assets processed.
Medical Research
3,000 research papers enriched, yielding hundreds of thousands of structured data points.
Education Publishing
Unified metadata and curriculum-mapped archive across the full primary-school content library. 100% curriculum coverage.
Admin, client, and content creator interfaces for indexing, projects, and asset management.
Ingestion, indexing, categorisation, and delivery APIs across the pipeline.
Automated topic labelling, relevancy cross-labelling, and content classification at scale.
Extracts people, dates, and topics, and supports processing across 140+ languages.
Separates text, images, tables, and links out of PDFs, scans, and legacy formats.
Orchestrated, resumable processing of very large archive batches.
Structured metadata storage with fast search across millions of assets.
Enriched data delivered into editorial, archival, and analytics tools already in use.
Elastic compute for large batch processing, with client-controlled data residency.
Separates Druid Learning, client, and user permissions and protects proprietary content end to end.
The team took the time to understand a genuinely complex problem before writing a line of code, and that showed in what they delivered. They built the indexing and categorisation engine our product depends on, stayed with us through several phases of development, and handed over clean documentation and full ownership at every stage. They have been a dependable technical partner as our business has grown.
Share a few details about what you're building, and we'll take it from there.