A dual-sided platform. For learners — a full study path without a tutor: AI dialogues with live corrections, games, vocabulary and speaking practice 24/7. For tutors — AI that builds assignments from each student's real mistakes, plus a live view of their practice. Below — the full process, stage by stage, with a real artifact for each. 63 screens are live: click through them.
Eine zweiseitige Plattform. Für Lernende — ein kompletter Lernpfad ohne Tutor: AI-Dialoge mit Live-Korrekturen, Spiele, Vokabeln und Sprechpraxis 24/7. Für Tutor:innen — AI, die Aufgaben aus den echten Fehlern der Lernenden baut, plus Live-Einblick in deren Praxis. Unten: der komplette Prozess, Etappe für Etappe, mit echtem Artefakt zu jeder. 63 Screens sind live — klick dich durch.

One senior designer, an AI-first workflow, a shipped product. I did the entire cycle myself — hypothesis-driven research (56 interviews), a two-role platform architecture, and 63 working HTML screens with a clickable wireflow — using AI at every stage: research synthesis, architecture, UI generation, copy, design QA and code. What a team does in months, delivered solo in weeks per release.
Eine Senior-Designerin, ein AI-First-Workflow, ein gelauchtes Produkt. Den kompletten Zyklus habe ich selbst gemacht — hypothesengetriebene Research (56 Interviews), eine Zwei-Rollen-Architektur und 63 funktionierende HTML-Screens mit klickbarem Wireflow — mit AI in jeder Etappe: Research-Synthese, Architektur, UI-Generierung, Texte, Design-QA und Code. Was ein Team in Monaten schafft — solo, in Wochen pro Release.
Before any pixels: three falsifiable hypotheses and a kill-criterion for each. Design started as a research question, not an opinion.
Vor jedem Pixel: drei falsifizierbare Hypothesen mit je einem Abbruchkriterium. Design begann als Forschungsfrage, nicht als Meinung.
| # | Hypothesis | Kill criterion | Verdict |
|---|---|---|---|
| H1 | Learners plateau because they don't speak, not because they lack grammar content | <50% name speaking practice as top obstacle | Confirmed — 82% |
| H2 | Fear of judgment blocks practice more than price does | Price ranks above judgment in survey | Confirmed — 67% fear judgment |
| H3 | Teachers will pay to see student AI practice, not to replace themselves | Teachers perceive AI as a threat | Partly — needs "teacher stays in control" framing |
Desk research at AI speed. I ran structured LLM scans over app-store reviews of 12 language apps, r/languagelearning threads and DACH market reports — then read the primary sources for everything that made it into a hypothesis. AI widened the funnel; judgment stayed manual.
Desk Research in AI-Geschwindigkeit. Strukturierte LLM-Scans über App-Store-Reviews von 12 Sprach-Apps, r/languagelearning-Threads und DACH-Marktreports — die Primärquellen zu allem, was in eine Hypothese einging, habe ich selbst gelesen. AI verbreiterte den Trichter; das Urteil blieb manuell.
Why kill-criteria matter: a hypothesis you can't lose is an opinion. Each one had a number that would have stopped the project — that's what made stage 02 research honest instead of confirmatory.
Warum Abbruchkriterien zählen: Eine Hypothese, die man nicht verlieren kann, ist eine Meinung. Jede hatte eine Zahl, die das Projekt gestoppt hätte — das machte die Research in Etappe 02 ehrlich statt bestätigend.
34 learners (A1–C1) and 22 teachers, plus a bilingual survey with 240+ responses across 18 countries. Learners speak their target language just 2.3 hours per month; teachers lose 8–12 hours a week to admin.
34 Lernende (A1–C1) und 22 Lehrkräfte, dazu eine zweisprachige Umfrage mit 240+ Antworten aus 18 Ländern. Lernende sprechen ihre Zielsprache nur 2,3 Stunden pro Monat; Lehrkräfte verlieren 8–12 Stunden pro Woche an Verwaltung.
Format. 30–40 min semi-structured interviews, recorded with consent; screening survey first (level, goals, current tools). Teachers recruited separately to avoid learner-teacher framing bias.
Format. 30–40-minütige semistrukturierte Interviews, mit Einwilligung aufgezeichnet; vorab Screening-Umfrage (Level, Ziele, aktuelle Tools). Lehrkräfte separat rekrutiert, um Framing-Bias zu vermeiden.
AI-assisted coding. Transcripts were clustered by LLM into a 14-node theme tree; I verified every cluster against the raw quotes and merged it down to 4 load-bearing themes. Rule: no quote enters an artifact without me hearing the original audio.
AI-gestütztes Coding. Transkripte per LLM in einen 14-Knoten-Themenbaum geclustert; jeden Cluster habe ich gegen die Rohzitate verifiziert und auf 4 tragende Themen verdichtet. Regel: Kein Zitat kommt in ein Artefakt, ohne dass ich das Original-Audio gehört habe.
Quant validation. The 240-response bilingual survey (18 countries) turned themes into numbers: 82% speaking-practice gap, 67% fear of judgment, 2.3 h/month actual speaking time.
Quantitative Validierung. Die zweisprachige Umfrage mit 240 Antworten (18 Länder) machte aus Themen Zahlen: 82% Sprechpraxis-Lücke, 67% Angst vor Bewertung, 2,3 h/Monat tatsächliche Sprechzeit.
| Aware | Try | First dialogue | Habit | Plateau risk | |
|---|---|---|---|---|---|
| Doing | Googles "speak German app" | Placement test | First AI chat | Daily 10-min sessions | Skips 3 days |
| Feeling | 😐 skeptical | 🙂 curious | 😰 → 😄 relief | 😄 confident | 😟 guilt |
| Pain | "another Duolingo?" | fear of being tested | fear of judgment | correction fatigue | streak anxiety |
| Design response | landing shows a real dialogue | placement framed as conversation | AI never says "wrong" — shows why | 3 correction modes | streak-repair, not shame |
| Aware | Try | First value | Habit | |
|---|---|---|---|---|
| Doing | Student mentions "an AI app" | Opens live view of a session | Sees exactly which grammar broke | Assigns topics between lessons |
| Feeling | 😠 threatened | 🤨 testing it | 😮 "this saves my prep" | 🙂 in control |
| Design response | "AI practices, you teach" messaging | read-only by default | error digest per student | "steer the AI" panel |
Full map: 5 personas (3 learner archetypes by CEFR + motivation, 2 teacher archetypes by tech attitude) × 6 stages, built in two workshops on top of the AI-drafted skeleton.
Vollständige Map: 5 Personas (3 Lern-Archetypen nach CEFR + Motivation, 2 Lehr-Archetypen nach Tech-Haltung) × 6 Etappen, in zwei Workshops auf dem AI-Skelett aufgebaut.
12+ platforms audited. The gap was clear: nobody combined judgment-free speaking practice with a teacher who stays in the loop.
12+ Plattformen auditiert. Die Lücke war klar: Niemand kombinierte urteilsfreie Sprechpraxis mit einer Lehrkraft, die eingebunden bleibt.
| Speaking 24/7 | Explains corrections | Teacher in the loop | Adaptive level | Gamified retention | |
|---|---|---|---|---|---|
| Duolingo | — | — | — | partly | strong |
| italki | scheduled | human | is the teacher | yes | — |
| Babbel | scripted | rules only | — | partly | basic |
| Speak/Talkpal | yes | opaque | — | yes | basic |
| Verbly | yes | shows reasoning | live view + steer | CEFR-adaptive | league + streak-repair |
Beyond features: pricing models (subscription vs credits vs seats), onboarding length in taps, how each product handles the first 60 seconds, paywall placement, and tone of error messaging. Two findings shaped Verbly directly: every competitor interrupts mid-exercise with its paywall (we never do), and none of the AI-speaking apps explains why a correction is right — the single biggest trust gap from stage 02.
Über Features hinaus: Preismodelle (Abo vs. Credits vs. Seats), Onboarding-Länge in Taps, die ersten 60 Sekunden jedes Produkts, Paywall-Platzierung und Tonalität der Fehlermeldungen. Zwei Befunde prägten Verbly direkt: Jeder Wettbewerber unterbricht mitten in der Übung mit der Paywall (wir nie), und keine AI-Speaking-App erklärt, warum eine Korrektur stimmt — die größte Vertrauenslücke aus Etappe 02.
| Must (MVP) | Should (v1.1) | Won't (deliberately) |
|---|---|---|
| AI dialogue with visible reasoning · placement-as-conversation · vocabulary loop (flashcards, quiz) · paywall & billing · teacher live view | League & achievements · listening drills · custom topics · teacher assignments | Video calls · human tutor marketplace · content authoring — the competitors' game, not ours |
Architecture: 2 roles (learner / teacher) × 5 zones (learn · practice · progress · monetization · account) — max 2 levels deep anywhere.Architektur: 2 Rollen (Lernende / Lehrkraft) × 5 Zonen (Lernen · Üben · Fortschritt · Monetarisierung · Konto) — überall max. 2 Ebenen tief.
| Zone | Web (33) | Mobile (30) |
|---|---|---|
| Learn | landing · signup · onboarding · placement · dashboard · topics · custom topic · AI dialogue | welcome · goal · placement · home · topics · custom topic · chat |
| Practice | crossword · match · builder · flashcards · quiz · session complete | crossword · match · builder · flashcards · quiz · listening · complete |
| Progress | vocabulary · progress · league · achievements · profile | vocabulary · stats · league · achievements · profile |
| Monetization | pricing · checkout · success · failed · billing · cancel flow · out of minutes | paywall · checkout · success · out of minutes · billing |
| Trust & edge | streak lost · empty/error states · role select | streak lost · offline |
| Teacher | dashboard · student detail · live chat view · assign | home · chat view · assign · assignment |
How AI helped: the first map had 94 screens. Three AI-assisted passes against the user stories cut 31 of them — every cut argued in writing ("which job does this screen serve?"). Cutting is the senior part; generating was the cheap part.
Wie AI half: Die erste Map hatte 94 Screens. Drei AI-gestützte Durchgänge gegen die User Stories strichen 31 davon — jede Streichung schriftlich begründet („welchen Job erfüllt dieser Screen?"). Das Streichen ist der Senior-Teil; das Generieren war der billige.
Every scenario was planned as a screen-to-screen map — and then built as a clickable map of live screens, where the student and teacher versions of the same moment sit side by side.
Jedes Szenario wurde als Screen-zu-Screen-Karte geplant — und dann als klickbare Karte aus Live-Screens gebaut, in der Schüler- und Lehrer-Sicht desselben Moments nebeneinanderstehen.











Every flow has its unhappy branch designed to the same fidelity — failed payment, cancel flow, out of minutes. In the full wireflow, student and teacher versions of each scenario sit side by side, web next to mobile.
Jeder Flow hat seinen Unhappy-Branch in gleicher Qualität — fehlgeschlagene Zahlung, Kündigungs-Flow, Minuten aufgebraucht. Im vollen Wireflow stehen Schüler- und Lehrer-Version jedes Szenarios nebeneinander, Web neben Mobile.
The visual bet: calm, not gamified-loud. Indigo as the "focus" color, warm paper background, one accent for teacher context. Type, radius and spacing tokens defined once — 63 screens stayed consistent.
Die visuelle Wette: ruhig statt laut-gamifiziert. Indigo als „Fokus"-Farbe, warmer Papierhintergrund, ein Akzent für den Lehrkraft-Kontext. Typo-, Radius- und Spacing-Tokens einmal definiert — 63 Screens blieben konsistent.
Semantic roles, not decorative colors: indigo = your learning, amber = teacher presence, green/red reserved strictly for correctness — so color itself teaches.Semantische Rollen statt Deko-Farben: Indigo = dein Lernen, Amber = Lehrkraft, Grün/Rot strikt für Korrektheit — die Farbe selbst lehrt.
| Token group | Decision | Why |
|---|---|---|
| Type | Display for numbers/headings, humanist sans for dialogue text | dialogue must read like conversation, not UI |
| Radii | 3 steps only (8 / 14 / full) | friendly without becoming toy-like |
| Spacing | 4px base grid, 8 named steps | 63 screens, zero ad-hoc margins |
| Components | ~40: chat bubbles with reasoning slot, correction chips, XP bar, streak calendar, league rows, exercise cards… | every exercise type composes from the same kit |
| Dark surfaces | reserved for "moment" screens (streak lost, offline) | emotional contrast where it matters |
AI's role: three moodboard directions and the first token draft were AI-generated; I picked, tightened and enforced. The kit lives as CSS custom properties in the shipped code — the tokens above are read from the product, not from a slide.
Rolle der AI: Drei Moodboard-Richtungen und der erste Token-Entwurf waren AI-generiert; ausgewählt, geschärft und durchgesetzt habe ich. Das Kit lebt als CSS-Custom-Properties im ausgelieferten Code — die Tokens oben sind aus dem Produkt gelesen, nicht von einer Folie.
| Situation | Typical app | Verbly |
|---|---|---|
| Grammar mistake | "Wrong. The answer is X." | "Almost! Word order flips after 'weil' — here's why …" |
| Lost streak | "Your 14-day streak is gone!" | "Life happens. One session today repairs your streak." |
| Out of minutes | Hard paywall mid-sentence | Finish the sentence first — then the offer |
Instead of a Figma clickthrough, the whole product was prototyped as 63 working HTML screens with real states, real copy and a linked wireflow. Engineers estimated against reality; tests ran on behavior, not pictures.
Statt eines Figma-Klick-Dummys wurde das ganze Produkt als 63 funktionierende HTML-Screens prototypisiert — echte States, echte Texte, verlinkter Wireflow. Entwickler schätzten gegen die Realität; Tests liefen auf Verhalten, nicht auf Bildern.




Why code instead of clickthroughs: engineers estimate against real markup, usability tests run on real behavior (scroll, states, focus), and stakeholders click a product — not a picture of one.
Warum Code statt Klick-Dummys: Entwickler schätzen gegen echtes Markup, Tests laufen auf echtem Verhalten (Scroll, States, Fokus), und Stakeholder klicken ein Produkt — nicht dessen Abbild.
8 moderated rounds (10–15 participants each) plus unmoderated sessions. Sources: moderated usability tests and product analytics on the beta cohort.
8 moderierte Runden (je 10–15 Teilnehmende) plus unmoderierte Sessions. Quellen: moderierte Usability-Tests und Produkt-Analytics der Beta-Kohorte.
| Finding | Change | Effect |
|---|---|---|
| Corrections mid-sentence broke flow | 3 correction modes (aggressive / balanced / minimal) | timing satisfaction 62% → 91% |
| Grammar explanations too academic | CEFR-adaptive "explain like I'm five" | comprehension 47% → 88% |
| Lesson builder abandoned halfway | Guided wizard instead of blank canvas | completion 44% → 87% |
| Pronunciation scores felt like grades | Reframed as "similarity to a native speaker" | adoption 59% → 84% |
Cadence. One moderated round per release cycle: 5 core tasks (first dialogue, change correction mode, finish an exercise loop, find progress, recover a streak), think-aloud, 10–15 participants mixed A2–B2. Between rounds: unmoderated first-click and comprehension tests on the live HTML prototype.
Kadenz. Eine moderierte Runde pro Release-Zyklus: 5 Kern-Tasks (erster Dialog, Korrekturmodus wechseln, Übungsloop abschließen, Fortschritt finden, Streak reparieren), Think-aloud, 10–15 Teilnehmende A2–B2. Zwischen den Runden: unmoderierte First-Click- und Verständnistests am Live-HTML-Prototyp.
AI in the loop. Session recordings were summarized per task by AI before I watched them — I reviewed flagged moments first, full sessions second. Before each round, an AI heuristic audit (Nielsen + platform patterns) caught the cheap problems so participants' time went to the expensive ones.
AI im Loop. Session-Aufnahmen wurden vor dem Ansehen per AI pro Task zusammengefasst — markierte Momente zuerst, volle Sessions danach. Vor jeder Runde fing ein AI-Heuristik-Audit (Nielsen + Plattform-Patterns) die billigen Probleme ab, damit die Zeit der Teilnehmenden in die teuren floss.
Mobile is not the web squeezed: thumb-zone actions, bottom navigation, native-feeling exercise interactions, offline state as a first-class screen.
Mobile ist nicht das gequetschte Web: Thumb-Zone-Aktionen, Bottom-Navigation, nativ wirkende Übungsinteraktionen, Offline-State als vollwertiger Screen.





The placement test is disguised as the first conversation: no exam anxiety, value before signup, and the paywall never interrupts mid-sentence.
Der Einstufungstest ist als erstes Gespräch getarnt: keine Prüfungsangst, Wert vor der Registrierung, und die Paywall unterbricht nie mitten im Satz.



Retention mechanics that respect adults: a weekly league, achievements tied to real skills — and a streak system that offers repair instead of shame.
Retention-Mechaniken, die Erwachsene respektieren: Wochen-Liga, Achievements für echte Skills — und ein Streak-System, das Reparatur statt Scham anbietet.



The unglamorous screens got the same care as the hero flow — because churn lives here, not on the dashboard.
Die unglamourösen Screens bekamen dieselbe Sorgfalt wie der Hero-Flow — denn Churn wohnt hier, nicht auf dem Dashboard.



