SUBTXT · AI Product Designer · 2026

Designing an AI-assisted contextual captioning tool

I teamed up with two developers to design and ship SUBTXT — a product that explains concepts in video conversations that casual audiences miss.
expressed through a system blueprint and exploratory interface designs.

ACHIEVEMENTS

  • Identified a real-world use case for an emerging AI capability
  • Shaped how the AI decides what to  explain flag, skip, and explain
  • Designed creator-controlled workflow
  • Prototyped in code to refine complex interactions

Role

Strategy, Design, Branding

Team

Me, Developer, Developer

Me, Founders (CTO/PM)

Year

2026

No viewer left behind.

Identifying concepts your audience may not know
Detecting references speakers don't explain
Prioritising moments worth captioning

ACHIEVEMENTS

THE IDEADEVELOPMENT

A system that captions video automatically

SUBTXT began as a 24-hour hackathon with two engineers. While exploring how Hera’s generative video API could work alongside Google Gemini, we realised we had the basis of a system that could automatically caption unexplained references and concepts, helping expert conversations reach wider audiences.

3 research artifacts available on desktop

🔍 What Focus Mode's design revealed

Focus Mode's design suggested a narrow framing of the problem - treating distraciton as something external, to be managed by controlling the environment.

🧠 A behavioural lens

To better understand focus and distraction, I turned to behavioural frameworks. In Indistractable, Nir Eyal writes:

While we love to blame external triggers...most of our distractions begin from within

This reinforced my suspicion that we were conceiving the problem too narrowly and provided inspiration for concept testing.

📈 The commercial opportunity

This misalignment wasn't just theoretical, it was reflected in the tools available.

🗺️ The product hypothesis
🗺️ Why Hera was perfect for creating captions
🗺️ The Problem

Expert conversations often contain references casual viewers do not understand, which go unexplained by the hosts.

⚠️ The Solution

A product that automatically detects these moments and generates short, contextual captions to explain them.

What was said

What SUBTXT captions

AI PRODUCT DESIGNDEVELOPMENT

Designing the creator experience

LLMs and video-generation tools made identifying references and generating contextual captions relatively straightforward. The harder challenge was designing the creator experience around that automation. I defined how creators should review and shape the AI’s work, then turned those interactions into a step-by-step workflow.

🤝

AI leads, creators decide

🪄

Make it easy for non-editors

💎

Make value visible at every step

The Workflow

Instead of exposing every decision in a Premiere Pro-style canvas, I separated them into focused steps that non-editors could work through progressively.

Division of labour

Rather than a one-shot handoff, I introduced review points into the automated process: AI recommends and generates; creators approve, adjust or override.

The workflow gives creators clear points to review and shape the AI’s work, without the complexity of a traditional editing canvas.

INTERACTION DESIGNDEVELOPMENT

Prototyping the ‘Curate’ interaction

The Curate step looked straightforward on paper, but prototyping it in code exposed tensions between recommendations, creator decisions and caption timing. The video below shows how the interaction evolved as I worked through them.

Video coming soon

Users can browse and interrogate moments

System recommends captions to includeents

Moments that clash are flagged

Requirements brief for Claude

Early prototype

Early prototype

PROTOTYING IN CODEDEVELOPMENT

Designing the "Curate" step

The Curate step looked straightforward on paper, but in practice AI recommendations, creator decisions, and caption timing affected one another in unexpected ways. Rapid prototypes with Claude made these interactions visible early.

The demo landing page communicated the concept through a working example.

Caption settings gave creators control over the AI's output.

Users can browse and interrogate moments

System recommends captions to includeents

Moments that clash are flagged

Requirements brief for Claude

Early prototype

Early prototype

Key Design Decisions


This walkthrough above focuses on what users experience. The notes below explain the reasoning behind the most important design decisions. (Coming soon!)

🗺️ 1. Defining the use case

🔒

🔒 Coming soon

🗺️ 2. Separating curation from editing

🔒 Coming soon

🔒

🗺️ 3. Balancing the experience

🔒 Coming soon

🔒

⚠️ 4. Finding the right interaction patterns

🔒 Coming soon

🔒

  • Users had already expressed interest in accountability-based support
  • It addressed a clear gap in existing focus tools
  • It aligned with behavioural research on how focus actually breaks down
  • Early thinking suggested the technical implementation was feasible

Key Screens

The demo landing page communicated the concept through a working example.

Caption settings gave creators control over the AI's output.

The Redesigned Experience

The prototype proved the concept, but it still reflected a demo rather than a product. I continued developing the concept independently, prioritising user control and clarity.

  • Users had already expressed interest in accountability-based support
  • It addressed a clear gap in existing focus tools
  • It aligned with behavioural research on how focus actually breaks down
  • Early thinking suggested the technical implementation was feasible

LANDING PAGE

Designing the AI's behaviour

The product relied on a multi-step pipeline: Gemini identified unfamiliar references before generating contextual captions, which Hera rendered directly into the video. I designed the editorial, writing and presentation rules that guided each stage of that process.

Video in

Flag moments worth explaining

Generate multiple caption briefs

Render captions, consistently

?

?

Video out

What we shipped in 24 hours

With limited time to design the interface, I focused on the two capabilities that best demonstrated the product's value: explaining unfamiliar concepts and giving creators control over the AI's output.

The demo landing page communicated the concept through a working example.

Caption settings gave creators control over the AI's output.

The demo landing page communicated the concept through a working example.

Caption settings gave creators control over the AI's output.

Evolving the interaction model

After the event, I continued developing the concept independently. While the prototype demonstrated the technology, it lacked the level of control needed for a professional creative workflow. Rather than asking creators to trust a one-shot AI output, I redesigned the experience around a series of editorial decisions.

SUBTXT — Four steps
STEP 01
⬆️ Upload
Get the video in, and confirm it's the right one.
  • Provide the video
  • Confirm it's the right one
  • System reads market, language, lengthAI
STEP 02
✍️ Curate
Decide which moments make the cut, at a sane volume and rhythm.
  • See everything found
  • Understand each reference
  • Understand why it was flagged
  • Include or exclude
  • Judge total volume
  • Judge cadence
  • Preview in context
STEP 03
🎬 Edit
Refine each caption — copy, imagery, timing and position.
  • Edit title and explanation
  • Change icon
  • Swap or remove image
  • Adjust timing and duration
  • Set on-screen position
  • Space out flagged clusters
STEP 04
⬇️ Export
Apply brand, watch it through, deliver the finished video.
  • Apply brand font and colours
  • Final watch-through
  • Export and download

A collaborative interaction model replaces the one-shot approach from the demo, offering a better experinece with more confidence and the ability to control alll the key asepcts of the product's output.