linkedin ads
img
druid learning

Industry Served: EdTech

Druid Learning, Turning Decades of Content Archives Into AI-Ready Data

An End-to-End Content Intelligence Platform for Publishers, Broadcasters, and Research Institutions


From Content Mess to Usable AI Data

Discuss your project

Druid Learning is a Dublin-founded technology company that transforms decades of archives, content libraries, and information repositories into AI-ready data. Publishers, broadcasters, and research institutions sit on enormous proprietary content estates: backlist titles, scanned books, audio-visual files, legacy flash material, PDFs, and metadata scattered across a dozen systems. Very few have the infrastructure to safely optimise and capitalise on it.

We developed the entire platform: a data intelligence layer that ingests content from any source, uses machine learning and natural language processing to enrich and standardise its metadata, and activates the resulting proprietary dataset inside the customer's existing analytics, editorial, and AI applications. Crucially, the source content is never altered and data ownership never leaves the client.

The platform now serves organisations including NPR, CJ Fallon, the Dutch Cancer Institute, and the University of Bristol, and is backed by Techstars, Enterprise Ireland, the European Union's Horizon programme, and LvlUp Ventures. Druid Learning was named winner of Frankfurter Buchmesse's 2026 Wildcard.

img
img

Project Challenges & Needs

Fragmented, Multi Format Archives

Content sat on FTP servers, in legacy databases, and behind APIs, as OCR, XML, INDD, MP4, PDF, EPUB, and legacy flash files.

Manual Micro-Categorisation

Itemising blocks of digital content and tagging component assets was manual, slow, and expensive to manage at scale.

Inherited, Inconsistent Metadata

Libraries acquired through acquisition or built on author-submitted records had no consistent, usable schema.

No Way to Break Files Apart

A single PDF could hold images, text, tables, and links with no means of separating them into individually addressable, reusable assets.

Our Solution

We built a scalable, modular Content Intelligence Platform enabling:

img
A three-tier user model covering Druid Learning, client organisations, and their content creators
Single and bulk file upload, with document name, type, page count, user, version, & upload date captured on ingest
Indexing that breaks each file into assets of the same type, separating images, text, tables, and links
Custom indexing across a chosen page range rather than the whole document
Indexing history, saved projects, and project creation from any indexed file
Asset folders defined by folder name, topic area, level or grade, and subject
AI categorisation of any asset, individually or in bulk, from the indexing output
Export of individual assets, folders, or whole projects to local storage
API delivery of the enriched dataset into analytics, CMS, and AI applications

The platform is delivered as connected products sitting on one pipeline, so clients can start by indexing and categorising an archive and expand into analytics, editorial suggestions, and metadata standardisation without changing their IT infrastructure or existing content.

Interfaces We Developed

Druid Learning Administrator Interface

The highest level of administration, allowing the Druid Learning team to:

  • Add, edit, delete, activate, and deactivate every client organisation on the platform
  • Monitor indexing and categorisation volume across all client accounts
  • Manage taxonomies, parameter layers, and platform-wide configuration

Client Administration Interface

A workspace for each publishing house, university, or institution, enabling client authorities to:

  • Add, edit, delete, activate, & deactivate their users
  • Define proprietary categories, taxonomies, and labels for their content estate
  • Review AI-generated tags and validate sample datasets before scaling up
  • Track saved projects with created date, indexed files, and total assets

Content Creator Interface (Indexing & Categorisation)

The core working environment, built around an Index menu covering Add File, Indexing History, Saved Projects, and Create New Project. Content creators can:

  • Upload files singly or in bulk, with type, page count, user, version, and date captured on ingest
  • Index a whole file, every uploaded file at once, or a chosen page range
  • Review output showing total assets and a breakdown by image, text, table, and link
  • Browse assets by type or view every asset from a file on one page
  • Send any asset through the AI categorisation flow
  • File assets into folders defined by topic area, level or grade, and subject
  • Build projects from indexed files, assets, and whole folders
  • Export any asset, folder, or project to local storage

Key Features & Capabilities

Centre & Course Setup

Admin defines courses, levels, sessions, and onboards each learning centre and DOS.

AI-Powered Micro-Categorisation

Machine learning labels every file by content topic, cross-labels for relevancy, and semantically links related content across the archive.

Custom Page-Range Indexing

Creators index a chosen range of pages rather than reprocessing entire documents, with full indexing history retained.

Projects, Folders, and Export

Assets organised into projects and folders by topic area, level or grade, and subject, then exported at asset, folder, or project level.

Human-in-the-Loop Validation

Sample datasets validated with subject-matter experts before scaling, built into the process rather than bolted on.

Client-Owned Data

Source content is never altered and clients retain full ownership and control of their dataset.

Outcomes and Impact

500k+

Assets Processed

100k+

Data Points Generated

140+

Languages Supported

100%

Of Your Data, Your Contro

img

NPR

Public Broadcasting

Audio, video, and editorial archives consolidated and transformed into a labelled, AI ready dataset. 500k+ assets processed.

Dutch Cancer Institute

Medical Research

3,000 research papers enriched, yielding hundreds of thousands of structured data points.

CJ Fallon

Education Publishing

Unified metadata and curriculum-mapped archive across the full primary-school content library. 100% curriculum coverage.

Technologies Used to Develop the Solution

Web Frontend

icon

Next.js

Admin, client, and content creator interfaces for indexing, projects, and asset management.

Backend

icon

Python (FastAPI) & Node.js

Ingestion, indexing, categorisation, and delivery APIs across the pipeline.

AI, Machine Learning

icon

Transformer Models & Classifiers

Automated topic labelling, relevancy cross-labelling, and content classification at scale.

Natural Language Processing

icon

spaCy & Named Entity Recognition

Extracts people, dates, and topics, and supports processing across 140+ languages.

Content Extraction

icon

OCR & PDF Parsing

Separates text, images, tables, and links out of PDFs, scans, and legacy formats.

Data Pipeline

icon

Apache Airflow & Kafka

Orchestrated, resumable processing of very large archive batches.

Database & Search

icon

PostgreSQL & Elastic search

Structured metadata storage with fast search across millions of assets.

Integration

icon

GA4, Chart beat, Tableau, WordPress, Salesforce, Airtable

Enriched data delivered into editorial, archival, and analytics tools already in use.

Cloud Infrastructure

icon

AWS & Google Cloud

Elastic compute for large batch processing, with client-controlled data residency.

Security

icon

Role-Based Access & Encryption

Separates Druid Learning, client, and user permissions and protects proprietary content end to end.

Here's What Our Client Has to Say

The team took the time to understand a genuinely complex problem before writing a line of code, and that showed in what they delivered. They built the indexing and categorisation engine our product depends on, stayed with us through several phases of development, and handed over clean documentation and full ownership at every stage. They have been a dependable technical partner as our business has grown.

Get in touch

Share a few details about what you're building, and we'll take it from there.

icon We respond within 24 hours.
Your data stays private and GDPR-compliant.

Marah Curtin

“Frekkel's AI turned our vision into something real. Our users love the personalisation, the transparency, and the progress they can see and feel.”

Marah Curtin (Founder & CEO, Frekkel)

Wendy Oke

“They're not just an outsource development company, they became an extension of our team.”

Wendy Oke (CEO, TeachKloud)

Dr Jake Robinson

“Square Root Solutions understood exactly what we needed for our AI-powered learning platform. The team delivered something our students and educators genuinely rely on.”

Dr Jake Robinson
(Founder & CEO, OnWard Education)

Cathal D’Arcy

“Square Root Solutions put me at ease from day one, clear communication, a brilliant project manager, and a team that treated our success as their own.”

Cathal D’Arcy (Founder & CEO, Bergo)

Deirdre Lyons

“Square Root Solutions brought both technical depth and genuine care to our platform. What they built gives our team real confidence in how it performs.”

Deirdre Lyons (Founder & CEO, Acuru)

Aaron Keane

“Square Root Solutions built a platform our students actually use every day. Reliable, easy to use, and exactly what we needed.”

Aaron Keane (Founder, Leaving Cert Plus)