Ideas · 22 entries

Ideas: disposable UI, malleable software, and learning by seeing

For forty years a UI was durable and costly: a team built it once and many people used it. In 2025 and 2026 it became something a model writes in a minute for one person and one question. This cluster covers the responses, which split into eight paths. (1) Malleable and home-cooked software (Ink & Switch, Sloan, Appleton) says regenerating apps is not the goal. People need to adapt and compose durable tools over shared data, and "AI code generation alone does not address all the barriers". (2) Per-prompt generated UI (Google Generative UI, Claude inline visuals): the labs' default is HTML written per answer and thrown away. Google's raters preferred it to markdown 83 to 91% of the time, with a wait of a minute or more hidden from them. Anthropic now draws a line in its product between permanent Artifacts and "temporary" visuals built "to aid understanding". (3) Catalogues (ChatGPT's roughly 70 concept modules, which is our inference; NotebookLM's fixed formats; Gemini's retrieved diagrams and quizzes): where correctness matters, labs pick fixed formats that the model fills in. (4) Instruments for understanding (Geoffrey Litt): when agents write the code, understanding is the bottleneck, and HUDs beat copilots. His examples are a one-minute custom debugger, a migration command center, and an explainer with a quiz. (5) The explorable-explanation craft (Victor, Case, Distill, manim) was right but too costly to author: Distill stopped in 2021 after editors spent 50+ hours on single articles. Agents now cut that cost, though generated visuals still have layout errors. (6) Reactive notebooks driven by agents (Observable, marimo pair). marimo pair puts the agent inside a live kernel where it reads values back exactly, which is the closest loop to fictty's. (7) Tutors, mostly in text. The evidence says behaviour design matters more than the model or the medium: Bastani found unguarded help harms learning, Kestin found a guarded tutor can beat a class, and Khanmigo's two-year trial shows 0.04 SD. Learning science backs relevant pictures plus words (g=0.37) and simulations for transfer. It warns that decorative visuals hurt and rejects "visual learners". No study tests a generated interactive screen against a text tutor. (8) The medium fight: Ptacek argues cheap UI means native apps, not TUIs. The replies point to remote use, Jane Street's finding that text snapshots make agents better at UI, and joshka's semantic-protocol idea, which Ptacek himself calls the strongest rebuttal.

The entries

terminal · harness · product

Claude learning mode and Claude Code Learning/Explanatory styles

Socratic mode in Claude, and Claude Code output styles that add insights or leave TODO(human) exercises in the code.

For fictty Ship a 'teach' skill pairing with these styles: walkthrough screens for insights, and a TODO(human) screen bound to live test status that waits via watch. Becomes a competitor if Anthropic brings visuals to Claude Code.

complementevidence: mixed
chat · renderer · product

Claude inline visuals

Claude builds temporary interactive HTML visuals inline mid-conversation, explicitly distinct from permanent Artifacts.

For fictty Adopt the durable-vs-temporary vocabulary; push screens unprompted when they help; compete only where HTML visuals can't reach: live data from commands, frame-rate interaction, exact read-back of what the person did, remote terminals.

competeevidence: strong
native · idea · essay

Stop Making TUIs (Thomas Ptacek)

Agents made native GUIs nearly free, so build CLIs and native apps, not TUIs.

For fictty Concede durable apps; his remote pattern (CLI on prod driven by a local UI) is fictty's socket architecture, so a local renderer of a remote screen answers him; keep the UI model semantic and independent of cells so it can move to a New Terminal or GUI.

competeevidence: strong
chat · renderer · product

ChatGPT study mode and interactive math/science visuals

Socratic study mode plus about 70 interactive concept visuals where you change variables and watch graphs update.

For fictty The catalogue path is what a lab chose where correctness matters most; fictty recipes (formula + parameters + plot bound to a computed source) are the terminal equivalent.

inspireevidence: thin
n/a · idea · essay

Malleable software (Ink & Switch)

Essay arguing people should adapt and compose their tools over shared data, not wait for vendors or regenerate apps.

For fictty Give people a gentle slope to tweak a pushed screen without regenerating it; keep data sources separate from screens so tools share data; use its scepticism of regenerate-everything as the case for a stable editable representation.

inspireevidence: strong
browser · idea · essay

On-demand instruments and AI HUDs (Geoffrey Litt)

Have the agent build a throwaway instrument (debugger, command center, explainer with quiz) so you understand, instead of asking it to fix things.

For fictty Make 'instrument' (step-through, scrubber, side-by-side) fictty's headline use case; quizzes and checkpoints need read-back via watch; expose history/seek as a UI component. Must beat 'ask for an HTML page' on live data, remote and read-back.

inspireevidence: mixed
terminal · idea · essay

The replies: TUI renaissance, Bonsai_term, semantic terminal protocol

The pro-terminal case: remote and keyboard use, Elm-style Bonsai_term whose text snapshots help agents, and a call for a semantics-not-cells protocol.

For fictty Cite Jane Street as outside evidence that text read-back makes agents better at UI; engage joshka with fictty's semantic UI value as a working example; document render --keys as snapshot tests for agent-pushed screens.

inspireevidence: mixed
any · renderer · library

3Blue1Brown and manim (and agents writing manim)

Python engine for programmatic math animation, now a target for agents generating explainer videos.

For fictty Layout errors dominate generated visuals; built-in layout in components avoids that bug class. Show transitions between bound states as steps (animate the data).

inspireevidence: strong
n/a · idea · essay

An app can be a home-cooked meal (Robin Sloan)

The founding essay of home-cooked software: a family messaging app with four users and zero churn.

For fictty Separate 'for few people' from 'for a short time' in our writing; provide a path from a pushed screen to a saved, reopened one, since some will be kept.

inspireevidence: strong
browser · idea · project

Distill

Peer-reviewed interactive ML journal that went on hiatus because interactive articles cost too much to make.

For fictty Use '50 hours per diagram vs one push' to state the cost collapse; quality needs a floor the runtime's components provide, since there are no editors.

inspireevidence: strong
n/a · idea · research

Dual coding, multimedia learning and the learning-styles myth

The evidence that words plus relevant pictures beat words alone, moderately, and that 'visual learners' is a myth.

For fictty Claim the multimedia principle, not learning styles; coherence (no decorative elements) and contiguity (callouts on the node) are design rules for generated screens; manipulable simulations have the better transfer evidence.

inspireevidence: strong
n/a · idea · research

Evidence on AI tutors (Bastani, Kestin, Oreopoulos)

Unguarded AI help harms learning, a well-designed tutor can beat a class, and real deployments show small effects.

For fictty Teaching screens must make the person act (predict, choose, answer) and never just show the solution; a small study of text tutor vs tutor-plus-screen on a developer task would be new evidence, cheap with exact read-back logging.

inspireevidence: strong
browser · idea · essay

Explorable Explanations and Dynamicland (Bret Victor)

The essay that named reactive documents and explorable examples: text as an environment to think in.

For fictty Add named parameters that bound sources depend on (the reactive document in fictty terms); teaching screens are guided walkthroughs with layers, not free dashboards; test value by whether the person can change an assumption and see the consequence.

inspireevidence: strong
n/a · idea · essay

Home-cooked software and barefoot developers (Maggie Appleton)

Talk arguing LLMs could let 'barefoot developers' build local community software, if someone supplies the glue.

For fictty The glue is the product: a runtime that draws, binds data and runs behaviour is what makes generated UIs usable. Keep the format plain data a non-programmer can read and edit.

inspireevidence: mixed
chat · idea · product

Khanmigo (Khan Academy)

Socratic text tutor inside Khan Academy exercises; a two-year RCT found 0.04 SD, about the same as practice without AI.

For fictty Engagement, not explanation quality, limited the best-funded tutor; screens should make the next action obvious and log exactly what the person did. Don't claim unmeasured learning gains.

inspireevidence: strong
browser · idea · research

Learn Your Way (Google Research)

Turns a textbook chapter into personalised text, narrated slides, mind maps and quizzes; beat a PDF reader on delayed recall in a small RCT.

For fictty Retrieval practice (quizzes with read-back) is the best-evidenced ingredient; one data source shown as several views is a natural fictty demo; generate screens from the real artefact, not model recollection.

inspireevidence: mixed
browser · idea · project

Nicky Case's explorables

Hand-built playable explanations of systems (Parable of the Polygons, The Evolution of Trust).

For fictty Predict-then-reveal with the guess read back (prompt node + watch); narration plus one control beats a dashboard of controls.

inspireevidence: mixed
browser · idea · product

NotebookLM (Google)

Source-grounded study tool that turns your documents into audio and video overviews, mind maps, quizzes and reports.

For fictty Every number on a teaching screen should trace to its source (show the command behind a bound value); a navigable tree-with-detail component is worth having.

inspireevidence: mixed
browser · idea · research

Generative UI (Google Research)

Gemini writes a custom interactive HTML page per prompt; raters prefer it to markdown when the wait is hidden.

For fictty Preference was measured with speed hidden and without comprehension; publish fictty's time-to-useful-screen and run a small pairwise study that adds a comprehension question. Keep the format small enough that valid means renders.

watchevidence: strong
browser · framework · library

marimo and marimo pair

Reactive Python notebook as a .py file; marimo pair lets Claude Code or Codex act inside the live kernel and read values back.

For fictty Keep agents acting on the live screen, with files as save/load; return per-node health after a patch; distribute as an installable agent skill.

watchevidence: strong
chat · idea · product

Gemini Guided Learning (Google)

Step-by-step Socratic mode in Gemini with retrieved diagrams, images, videos, quizzes and flashcards.

For fictty Quizzes are the cheap, evidenced part; prefer found visuals (existing diagrams, command-produced charts) over generated ones.

watchevidence: mixed
browser · framework · product

Observable (notebooks, Framework, Notebooks 2.0)

Reactive JavaScript notebooks with D3/Plot, now an open HTML file format plus a desktop app.

For fictty Dependency-driven updates from parameters are the missing piece for reactive teaching screens; if people edit saved screens, consider a friendlier authoring surface while keeping JSON canonical.

watchevidence: mixed

The overview

Ideas: disposable UI, malleable software and learning by seeing

The overview for the “ideas” cluster of the landscape (10 October 2026). It extends the first pass in the landscape post. Each entity has its own dossier beside this file.

The shift in one sentence

For forty years an interface was a durable, costly thing a team built once and many people used; in 2025 and 2026 it became something a model writes in a minute for one person and one question. Everyone in this cluster is working out what follows from that, and they disagree.

The paths people are going down

1. Make software adaptable, not disposable (malleable software, Robin Sloan, Maggie Appleton). The Ink & Switch camp says regenerating an app is not the goal; people need to change the tools they have, compose them over shared data, and make small edits without a conversation. “AI code generation alone does not address all the barriers to malleability.” Home-cooked software is about audience (a few people you know), not lifespan; Sloan’s family app has lasted six years.

2. Generate the interface per question (Google Generative UI, Claude inline visuals). The labs’ default: the model writes HTML for each answer and throws it away. Google’s study found raters preferred it to markdown 83 to 91% of the time, with the minute-long wait hidden. Anthropic made the distinction product policy: Artifacts are permanent, visuals are “temporary,” built “to aid users’ understanding.”

3. Fill a catalogue (ChatGPT visuals, NotebookLM, Gemini Guided Learning). Where correctness matters, the labs quietly pick fixed formats the model fills: about 70 interactive concept modules in ChatGPT (our reading), mind maps and slide decks in NotebookLM, retrieved diagrams and quizzes in Gemini. Reliable, narrower.

4. Build instruments for understanding (Geoffrey Litt). The sharpest version of the thesis: when agents write the code, the human’s bottleneck is understanding, and the fix is HUDs, not copilots: a one-minute custom debugger, a migration “command center,” an explainer with a quiz to pass before sharing. All throwaway, all bound to live program state.

5. Keep the old craft of explanation, now cheaper (Bret Victor, Nicky Case, Distill, manim). Explorable explanations were right and too expensive: Distill stopped in 2021 after editors spent 50-plus hours on single articles. Agents now write manim scenes and interactive figures, with layout errors common enough that research pipelines keep a human.

6. Durable reactive documents, now driven by agents (Observable, marimo). Notebooks are where data explanation already lives. marimo pair (2026) lets Claude Code or Codex act inside a live notebook kernel and read values back: the nearest thing in this cluster to fictty’s loop.

7. Tutors, mostly in text (Khanmigo, Claude learning mode, AI-tutor evidence, learning science). The evidence says behaviour design matters more than the model: unguarded help harms learning (Bastani, PNAS 2025), a well-designed tutor can beat a class (Kestin, 2025), and the largest deployment shows 0.04 SD because students don’t engage (Khanmigo, NBER 2026). The learning science supports relevant pictures plus words (g = 0.37 across Mayer’s studies) and simulations for transfer, warns that decorative visuals hurt, and rejects “visual learners.” No study we found tests a generated interactive screen against a text tutor.

8. The medium fight (Ptacek, the replies). Ptacek: cheap UI means build native, not TUI. The replies: remote and keyboard use, Jane Street’s finding that text snapshots make agents better at UI, and joshka’s call for a terminal protocol of semantics rather than cells. Ptacek himself names “the New Terminal” as the strongest rebuttal.

What’s really being solved

Under all eight paths is one problem: closing the gap between what the computer knows and what the person understands, at the moment they need it. Cheap UI matters because it lets the interface be specific to that gap. The debate is about who owns the interface’s lifecycle (the vendor, the user, or the model for a minute), how it’s checked (by a human designer, a schema, or nothing), and whether it changes understanding or only preference.

Three findings cut across the cluster:

  • Preference is not understanding. The headline numbers (Google’s 83%) measure liking. The learning evidence is modest and says the interface must make the person act (predict, choose, answer) and must avoid decoration.
  • Data and code split along reliability. Where correctness matters most (teaching maths, citing sources), labs choose catalogues and grounding over free generation.
  • Live state and read-back are the new requirement. Litt’s instruments bind to running programs; marimo moved agents into the live kernel; Jane Street’s agents check screens as text. The loop where agent and person share one live surface, and the agent reads it exactly, is where the field is heading.

Where fictty stands in this cluster

Honestly: in teaching and explanation, the browser wins on richness, and the labs (Claude visuals, Gemini, ChatGPT) own the general case. fictty shouldn’t pitch itself as an education product.

Its defensible ground here is narrower and real: instruments for understanding what an agent is doing, in the terminal where the agent works. Litt’s debugger and command center, built as data bound to live sources, readable back exactly, and with checkpoints (watch) where the person predicts or answers. That combines path 4’s purpose, path 3’s reliability, path 6’s live loop and the learning evidence’s demand that the person act.

What to do because of this cluster:

  1. Make “instrument” (step-through, scrubber, side-by-side) the headline use case and demo.
  2. Ship watch and a question or checkpoint component; a teaching screen must read back what the person did.
  3. Add named parameters that sources depend on (Victor’s reactive document, Observable’s dataflow).
  4. Return health after a patch (marimo): which nodes rendered, which sources failed.
  5. Ship a Claude Code skill that pairs with the Explanatory and Learning styles.
  6. Keep the model semantic and independent of cells, so it can move to a “New Terminal” or a local GUI renderer (Ptacek’s own remote pattern).
  7. In the blog post, claim the multimedia principle and the cost collapse (Distill’s 50 hours), not “people are visual learners”; and measure time-to-useful-screen against generate-a-page.

Entities dropped or added

All entities in the brief exist and are covered. Added: Google’s Learn Your Way (a controlled test of generated multi-representation material), ChatGPT study mode and visuals (the catalogue counterexample), Claude inline visuals (the clearest lab statement of temporary UI), the AI-tutor evidence (Bastani, Kestin, Oreopoulos), and the TUI replies (Jane Street’s Bonsai_term, joshka’s protocol). Split: Claude learning mode and Claude inline visuals are separate dossiers; Sloan, Appleton and Litt are separate dossiers.

What we couldn’t verify

  • Whether ChatGPT’s interactive visuals are prebuilt modules (inferred from the fixed concept list).
  • How Claude’s inline visuals are built beyond “HTML.”
  • The exact arXiv date of Google’s Generative UI paper (listing and ID disagree).
  • NotebookLM’s “Cinematic Video Overviews” (one secondary source).
  • Whether Jane Street’s Bonsai_term is open source.
  • Any new Nicky Case explorable or Dynamicland release since 2024.