top of page

Intelligent Metadata, Search Optimization & Interactive Design for Media Archives

ARCHIVAL AI by 

Logo-05.png

Generate
Auto-tags, transcripts & OCR from video, image, audio, text

  • Schema-aligned fields (CDWA, CCO, Dublin Core or custom)

  • Content clustering to reveal emergent categories

  • Easily transition between ontologies or merge existing ones

  • Built-in interpretability and visualization: our AI is not a block-box

Govern
Consistency & quality you can trust.

  • Rules engine enforces tone, terminology & placement

  • Ontologies/authority files applied across the archive

  • Automated evaluation vs. human baseline

  • Audit trail and versioned outputs for review

  • Built on open-source technology, runs completely on-premise if needed

Connect & Discover
Your archive doesn't exist in isolation, and it shouldn't feel like it does.

  • Semantic/keyword/image search across all content types

  • “More like this” via vector similarity clustering

  • Standardized searchability, even with uneven metadata quality

  • Curator-friendly facets and filters

  • Interactive timelines, visual maps, and thematic clusters for public audiences

  • Automatic cross-referencing with relevant public collections (museum databases, national archives, library catalogs)

The Story of Archival AI

Drowning in files

___or discovering what matters?

A collection no one can navigate is a collection no one can love.

Archival AI by onto transforms raw media into structured, searchable, and publicly engaging archives, and connects them to the wider world of knowledge they belong to. We extract tags, transcripts, and on-screen text; map them to your schema; enforce consistency at scale; and build interactive environments that link your collection outward to relevant external archives, databases, and institutional records.

Because the goal was never just findability. It is connection.

 

Industries We Serve

Marketing  & Sales Enablement

Reuse marketing and training media, and make it instantly sales-ready. Search by campaign, product, segment, region, speaker, or rights; enforce approved & latest versions and cut duplicate shoots.

Cultural Heritage & Artist Estates

Keep heritage discoverable, alive, and in context. Archival AI drafts schema-aligned records (CDWA/CCO/DC), applies authority files, enforces terminology, and clusters related works. We co-design interactive environments and connect your collection to relevant external archives, peer institutions, and linked open data sources.

Media & Publishing

Find the right clip in seconds, present it in ways that resonate, and surface the broader context it belongs to. Index speech, on-screen text, visuals, and music; filter by transcript status, language, rights, or topic; jump to timecodes. Link footage to external event records, referenced people, and related collections. 

Universities & Research Labs

Make lectures and research media truly searchable: ASR transcripts, OCR, topic facets, and named-entity linking for citations - with cross-references to external repositories, citation databases, and institutional archives- plus interactive learning environments for students and the public.

Brand & Enterprise

Make your company’s history usable. Search past events, campaigns, launches and talks; copy time-coded moments and visuals straight into decks and wikis. No more hunting across drives.

Our
Solutions

Strategy

We inventory your media, systems, and policies; map sources and risks; and identify relevant external archives worth connecting to. If holdings are still physical (tapes/film/photos), we scope digitization & ingestion as part of the plan. We align on goals that span storage, governance, experience, and institutional interoperability.

Execution

Co-design your data model, authority files, and rules. Define QA workflows and curatorial logic: thematic groupings, narrative arcs, and interactive access points. Identify and integrate external archive connections [linked data pipelines, cross-institutional references, and shared authority files] human-in-the-loop throughout.

Data Advisory

We deploy Archival AI, centralize and index all your media, connect relevant external archives, and build interactive front-ends where appropriate: visual explorers, public portals, exhibition-ready interfaces. We hand over an archive that is governed, connected, and genuinely engaging.

Previous Work

Our relevant work that have been completed with partners from various industries.

Screenshot 2025-10-06 at 18.56.07.png
Nokia Design Archive

⁠Nokia Design Archive is a publicly accessible digital portal developed as a research project at Aalto University’s Department of Design, built from materials donated by Microsoft Mobile Oy and Nokia designers to study the role of design within big organizations like Nokia. The source archive comprises 20,000+ entries and about 950GB of files spanning from the mid-1990s to 2017. Lu Chen as a member of the research project began prototyping interactive visualizations in 2023, continued the web application development in the summer of 2024. The portal was launched in January 2025 with support from Aalto University Communications. This project “remediates” archival material into accessible knowledge through interactive visualizations for a non-academic audience.

Nokia Design Archive & Aalto University

Screenshot 2025-10-06 at 18.48.09.png
Teoman Madra Archive

Teoman Madra Archive is an conservation and digitization effort launched in 2021 to rescue and organize the multimedia oeuvre of Teoman Madra, a pioneer of Turkish media art, focusing on works produced between the 1960s and 2000s. The team (Begüm Çelik & Selçuk Artut) first retrieved materials scattered across locations, stabilized fragile carriers, and began systematic digitization and cataloging of slides, negatives, VHS/Betamax/miniDV tapes, optical media, and drives. Descriptions follow museum-grade standards (CDWA/CCO) and include both Madra’s artworks and historically valuable documentary recordings, building a research-ready corpus. 

Sabancı University & Madra Family

Justify_DemoDay_Pitch.001.jpeg
Tuumailubotti

Tuumailubotti is an experimental conversational AI that explores how a chatbot might represent neurodiverse rhetoric rather than defaulting to neuronormative styles. Built on Finnish FinGPT-3 models to achieve native-level Finnish proficiency, it was evaluated with 31 participants, and its curated dataset has been released openly for further research. In short, Tuumailubotti blends HR support, design research and local-language AI to question how conversational systems can better include neurodivergent ways of communicating.

A Large Finnish Media Company

bottom of page