Maggie Appleton on Design Engineering: Paper, Jigs, and What Agents Still Can't See

Open on YouTube ↗
Overview

Maggie Appleton is a staff research engineer at GitHub Next, where she prototypes ways software engineers might work with AI agents. She has been an illustrator, the first designer at the AI startup Elicit, and a front-end engineer. In this conversation on the Pragmatic Engineer Podcast, she explains what a design engineer is and why designers need to understand engineering constraints. She describes how her design process has changed with agents, and why she still starts nearly everything in a paper notebook. Her position is that agents have made implementation and prototyping much faster, but the early, visual, judgment-heavy part of design remains a human job. She is not sure how she would feel if that changed.

30 min read
3:46

From cultural anthropology to front-end engineering

Appleton studied cultural anthropology at university and calls it her "one true love." She describes it as the study of human beings through participant observation: living closely with people to understand how they see the world. What drew her in was how much cultures differ, not just in food or housing, but in how they understand color or time. She grew up as an expat child who left London at six, and the field confirmed something she already felt: there is no fixed way to build a society, and the range of what is possible is wider than people assume.

When graduation came, the career options her professors described were a PhD and professorship, or working for the military. She went elsewhere. As a 90s kid she had learned HTML and CSS around age 12 or 13 through Neopets and MySpace. She had also done IT work for her university. So she started freelancing in web design to pay rent, and gravitated toward illustration because she loved drawing.

Her first exposure to product design came at a design and development shop in Prague that built apps for San Francisco startups, including early work for Tinder and Uber. She made illustrations, logos and branding there, and realized "there are people who design buttons and sidebars." She then spent about four years as an illustrator at Egghead, the developer education company, and later became its art director. She says that was her real front-end engineering education. To illustrate React components, useEffect, and JavaScript functions, she had to understand them, and she gradually moved into front-end engineering and visual design.

8:04

Elicit: a feature a week for a year

Around late 2021 or early 2022 she joined Elicit, an AI startup that uses language models to speed up scientific literature reviews. That work traditionally means academics reading thousands of papers and copying data into spreadsheets, which she calls a perfect use case for models. One founder had studied language models for about ten years and was convinced they would change science.

The team had six or seven people when she joined and stayed small. She was the only designer the whole time. There was a PM for a while, who left and was not replaced. Elicit had a popular free prototype, which gave the team a lot of data on searches, clicks, and A/B tests. She says they shipped a feature every week for a year, in a loop of design, build, ship, measure. Over a little more than two years she learned product design end to end, from user needs and domain understanding through front-end implementation. She describes the experience as wonderful and intense, and says she left wanting "to step back from startups for a minute."

10:33

What a designer actually does: nouns and verbs

Asked what a designer does, Appleton says it is not that different from engineering. It is problem solving with different materials. The process is the same: define and scope the problem, explore possible solutions and their trade-offs, prototype, and validate with users. The materials are space, size, weight, color, motion, prominence, and copy, for example whether a button's words tell the user what it will do.

In product design, she adds, you also work with the "nouns and verbs" of a product. For a sneaker store the nouns are easy: sneaker, cart, money. For something like AWS they get very hard. She has mostly worked on power-user tools for scientists and developers, where the nouns become abstract: the right container for a set of data, how one dataset connects to another, what to call a function that transforms data. Developer tools are especially hard because they are so malleable, meaning they could take many forms. A government form is a hard design problem but not complex in the same way.

Her current work on agentic tools shows the problem. Nobody yet knows the nouns of agents. There are chat sessions, plans, MCPs, skills. A designer could create a new primitive that connects several sessions to a pull request, and each choice of boundary changes how users think about the product. Then come the verbs. Editing, renaming, and deleting are standard CRUD, but can you fork an agent session, and what would that imply? For her, the hard part of product design is reducing all that possible complexity to a coherent, canonical set of nouns where users can point at something, know what it will do, and not be surprised by what happens when they click. She says this is "very easy to say" and extremely hard to do.

The host compares this to Kent Beck and Ward Cunningham keeping a thesaurus nearby while naming the constructs that became design patterns. Appleton agrees. In tech, she argues, product design is really software design, and software design is engineering. A designer who cares about colors, shadows and motion also has to understand the database shape, the APIs, and how data flows through components.

33:20

Who counts as a design engineer

The host pushes back with experience from Uber and Skype. There, designers worked in Sketch and later Figma, ran explorations and user testing, and handed visuals and animations to engineers without knowing how the database worked. Appleton says her view reflects her own background. Not every domain requires deep technical engagement. For government forms, a designer probably doesn't need to understand the database and is better off advocating for users through interviews and flow design. But for developer tools, or in a new field like AI where the back end largely decides what the product can do, designers have to engage with the technology. She sees a spectrum from business and market context through users and interface down to the back end. Designers can be valid anywhere on it. She likes the slice that is half engineering and half design.

A few years ago she tweeted about wanting to become "a full-blown design engineer" and collected names of people doing that work. She now thinks the role exists, but defines it more narrowly than much of X does. Many people who call themselves design engineers there are, in her view, very good micro-interaction designers: hover effects, loading transitions. She says that work is cool, but most of it can now be done by telling an agent without reading any code, so she sees it as advanced motion and visual design rather than design engineering, and doubts it is a full career.

For her, a design engineer is still a designer, caring about product nouns, verbs and visuals, who also works closely with engineers and usually implements much of the product directly, with a deep grasp of its technical architecture. The questions are things like: given the shape of the back-end data, what is possible in the interface? Given what the models can do and what custom skills the product has, how do you explain its capabilities to users? She has always done her own front-end work. Engineers she has worked with were usually thrilled, she says, because they hate CSS and could focus on sync and logic instead of whether a border radius is right. They rarely care about that, and she does.

36:15

Designers must understand their materials

Because she has always been the front-end person, Appleton says she hasn't felt the classic tension where a designer delivers a perfect Figma file and the engineer doesn't build it to spec. The host has. They recall designs with gradients that would have wrecked scroll performance on mobile, and designers focused on iOS while engineers worried about low-end Android devices.

Appleton blames the tools more than the designers. Sketch and Figma pixel mockups, she says, "have no relationship to the constraints of the medium." Any designer has to understand their materials. A table designer knows how oak behaves and how pine dents. A web designer who doesn't understand performance, data fetching, loading times, or race conditions will produce bad design solutions and bad working relationships. She hopes agents will help here. Designers can use them to move toward engineering and as tutors: "the engineers told me about this error state I've never heard of — explain why it would occur and help me build a state machine." She finds it hard to imagine things going well for designers who still want to make only pixel mockups.

Her toolkit: everything, plus paper

Her tools change constantly, partly because her job at GitHub Next is to figure out what GitHub should do next. She dogfoods internal tools and tries every agentic harness she can: Codex, Claude Code, Conductor, Amp, OpenCode, Pi. She currently loves Codex and calls OpenAI's desktop app beautiful and well thought through. She still uses paper and pen for early brainstorming, and Figma because she knows it well.

Once she has worked out what to build, which she says is most of the work, she hands it to an agent. Implementation isn't "solved," but it is good enough that if she specs something well and lists how the agent should verify its work, she can say "let me know when you've got a PR up." She rarely looks at code now beyond skimming PRs in review. She opens VS Code occasionally. She notes that GitHub Next builds prototypes to validate ideas, not production software, so its quality bar is much lower than github.com's.

21:50

Sketching a better way to plan with agents

Appleton showed her notebook on camera. Her current prototype starts from a belief that planning is a bottleneck with agents. Today planning usually means a long chat, often in a CLI, where the agent asks a series of multiple-choice questions, sometimes a hundred of them, for example with Matt Pocock's "grill me" skill. The host says they went through 36 questions and got annoyed around question 20 or 25, with no idea when it would stop. Appleton says people tire and can't make that many decisions that quickly. They also lack enough information for each one, so when an option is marked "recommended" they just start accepting it. Agents love to produce reams of text, which is a good output for agents but "not an ideal input for humans."

Her sketches explore a graphical interface where each decision gets its own small document or card. Simple decisions can stay multiple choice. Others need to show more. If the agent asks whether borders should be 10%, 12%, or 15% gray, it should show the options visually. If it asks how to structure the architecture, it should show an architecture diagram, data-flow diagram, or state machine. She is prototyping decision cards with embedded diagrams, prototypes and HTML, nested inside a larger plan document. The interface should be multiplayer so a team can plan together. Each decision should have a named human owner for an audit trail. The point is not to hunt people down, but so that six months later someone can open the card and see what information was available when the choice was made. On paper she is trying out shapes: an expanding stack of cards, one long linear document, or swiping like Tinder.

The notebook also has lighter projects, including a character for an ESP32. She credits Steve Ruiz with persuading everyone to buy these small Wi-Fi and Bluetooth devices with screens. She wants it to tell her the weather and how much agent capacity she has left before her usage limit resets.

25:44

Why pen and paper aren't going away

Her sketching habit comes from illustration, where you have to draw to work out composition. She trained in LA with film concept artists whose method she calls very technical: building things from 3D shapes and understanding how robot joints work so a drawn robot is believable. That background made UI sketching come naturally.

Her argument for paper is about speed and looseness. She could tell Claude Code, "here's an idea, it's like a stack of cards and it's an accordion," but drawing it by hand is faster and takes much less effort. Early on you need quick, loose feedback to find the shape of something before you can put into words what you want an agent to do. A sketch also stays on the desk, so the next day she remembers what she was working out. The sketches don't have to be visual designs; data diagrams work too. The key is that they externalize thoughts that aren't linguistic.

She says agents mostly take text as input. They can read images, and she does give them photos, but they are weak at spatial reasoning and visual design. They leave out spacing, size things wrongly, and overlap text: "They can't see." So she still does much early design without agents, and is skeptical of people on Twitter who claim to involve agents earlier.

The host compares this to engineers at a whiteboard, where drawing boxes and arrows makes people put their ideas into shared 2D space. They add that Miro caught on by bringing that to remote teams. Appleton connects this to how primitive agent interfaces still are, only about four years after ChatGPT launched in November 2022. She recalls that at Elicit they had been debating a chat interface when ChatGPT came out. Software design itself is a young field, she says, perhaps 60 years old. Agents live in a world of weights, skills and MCPs, and humans live in one of physicality, texture and light. The hard challenge is finding artifacts where the two can meet. She is clear that she doesn't think agents are conscious. She calls them a kind of intelligence that thinks and acts differently from humans. Her "eventual dream" is an agent that can look over her shoulder at her notebook and help move her ideas forward, but she thinks that is a long way off.

She has also started woodworking. Partly it's a joke that every software engineer eventually picks a physical hobby like ceramics, bread, or motorcycle repair. But she also bought an old London house that needs shelves and a repaired banister. She used Claude, ChatGPT, YouTube and Instagram as coaches, then signed up for a course, expecting parallels to software "in a more satisfying way where you actually touch the thing you make."

48:03

Jigs: building your own Figma

Appleton uses Figma to move past notebook sketches to medium fidelity: exact contrast, text sizes. She never takes it to high fidelity, because things look different in the browser (font rendering, responsiveness, breakpoints), so she moves to a browser prototype early. Now she points an agent at the Figma file to build it. For high-fidelity mockups she has run a loop overnight: take screenshots, compare with the Figma file, keep going until they match. By morning the interface is roughly as specced. She calls this grunt work ("we're building another sidebar"), where code quality and tech debt don't matter because the prototype may be thrown away.

Her favorite technique is the "jig." The term comes from woodworking, where a jig is a small device built for one specific job. In a live prototype she tells the agent which values she's unsure of, such as headline size, colors, or an animation curve. The agent adds sliders and color pickers so she can tune them live, then commit the final values. She calls it "build your own Figma as needed," not limited by Figma's features.

She demoed a prototype for a new GitHub Next website: an animated, spinning constellation of the team's past projects, with physics on the stars, hover scaling, and a scroll transition into the page. She began with a notebook sketch, then built jigs for background color, rotation speed, gravity intensity, star variation, resting logo size on scroll, and the strength of a glass effect. She showed how too much glass washes out the text and too little makes the effect invisible, so the right value is somewhere between. She says she wrote no code and didn't need to know the physics. She has a skill that tells the agent how to make these controls look decent. The one in the demo used DialKit. Static mockups of this in Figma didn't work, she says, because you can't feel the layout or scroll distance until it is live.

She links this to Bret Victor's 2010–2013 talks, including "Stop Drawing Dead Fish." His point was that programming isn't live: you edit variables in code, build, and look at the result elsewhere. He argued for direct, instant feedback on the artifact. In her view, jigs finally make that possible. Before, the closest thing was browser dev tools, and hand-building such controls wasn't practical because coding the thing itself was slow enough.

50:39

"Do we still need designers?"

The host quotes an engineer who asks Claude to generate 20 high-fidelity alternatives, iterates to a "perfect design," and finds it faster and more enjoyable than working with a design team. Appleton thinks that's fine when no designer is available and you need an interface to validate a hypothesis, especially for something simple like a button and a sidebar. Taste matters, though, and she might see those designs as obviously AI-generated and unlikely to do their job. When you are defining a new primitive, you need someone whose job is to think the problem through: run experiments, show them to users, and learn what people actually understand.

When models design for her, she judges the quality as poor. She thinks they've been prompted with universal design principles but lack nuance and context, so they apply the principles too rigidly. Her example is labels. Where a designer would use an icon button to close a sidebar, because users know the convention and will try clicking it, the agent writes "Close sidebar" or "Close modal," or adds four lines of instructions above a button. She understands the model is trying to make the interface explainable, "but actually it's really bad design."

53:42

The UI of AI may just be a table

Asked what good UX for AI looks like, Appleton says the question won't be answered for a while, but many principles haven't changed. At Elicit, scientists extracted paper data into Excel, one paper per row. The team thought that was old-fashioned and tried infinite canvases of scattered cards, linear card stacks, and Notion-like composable documents. In every interview users found them confusing and asked, "Can I just have a table?" After a couple of months the team concluded the new UI of AI was, at least for their case, a table. The familiar interface had the lowest cognitive load. They later built many features into and around the table.

She expects AI interfaces to keep using documents, sidebars, cards and tables, starting from familiar primitives and expanding outward. Chatbots weren't the final form. Codex still has a chat window but adds git worktrees, annotations, and debugging around it. The host adds that every new paradigm has a learning cost, like teaching parents or grandparents to use phones, and people don't want to pay it again.

58:40

Capability gaslighting

Appleton coined "capability gaslighting" before reading Ethan Mollick's similar "jagged frontier." Models are very good at some things and very bad at others, and it's hard to predict which a given task will be. Their impressive successes convince you they're highly capable, then they fail at something else. She says she sometimes misses how badly a model has failed because she still believes "Opus could never really get this wrong," and it "totally does all the time."

She contrasts this with human experts, who rarely lose their expertise suddenly. If they did, you'd wonder if they were having a breakdown. Models can do a task well one day and fail the next because of different prompting, different context, or plain stochastic variation. She grants that frontier models have a floor they won't drop below, but says upgrades can shift what a model is good at in ways that are hard to predict.

1:00:44

One developer, two dozen agents, zero alignment

Her talk with this title describes a problem GitHub Next has worked on for more than a year. Agents make each person much faster, but software is built by teams that need to agree with product managers, designers and other engineers. Slack, Linear, and GitHub Issues exist, but much planning happens before an issue is written, because by then you're usually ready to hand off to an agent. Teams need to agree on whether to build a feature, its shape, its interface, and any database migrations. She says there are few good tools for that stage, especially ones involving agents.

The biggest gap, she says, is that agentic coding tools aren't real-time multiplayer. Sessions are private and often local. That is starting to change. She mentions Buzz from Jack Dorsey and GitHub Next's own ACE prototype, which offered shared compute and sandboxes in a Slack-like interface so people could code and talk at once. But even with an agent in a chat channel, teams still need ways to agree, which is why she wants decisions to become a first-class primitive: large enough on screen to inform the choice, open to teammates' input, owned by someone, and recorded as context for the agent. Before, implementation took long enough that teams could adjust along the way. Now there is a hard handover to an agent, so alignment has to happen up front, and she says her own team struggles with this too.

ACE combined Slack-style chat with cloud microVM sandboxes, PR creation, and code review. She says it was a useful prototype but too ambitious for a team of three or four, and couldn't become a full product. Parts of it are being shipped. Sandboxes are now part of the GitHub desktop app. She thinks it also made leadership take multiplayer more seriously. The host describes Ramp's internal agent, Inspect. It works in Slack, a Chrome extension, and a web interface, and all sessions are public so anyone can join and redirect them. Ramp worried about privacy but found it wasn't a problem, and the openness helped adoption. The host credits its success to deep integration with Ramp's internal systems and full cloud developer machines, which is easier for one company than as a generic product.

GitHub Next's role, she explains, is to run ahead of the product teams and explore bigger or riskier bets. After multiplayer sessions, her team is now looking at proactive agents that aren't annoying. They start from an assumption that intelligence is nearly free. If you had 100 background agents, what is the least overwhelming way to consume their output? How could agents work in a shared document without feeling invasive? She calls these open research questions and hopes the prototypes will inform GitHub if it builds in this direction.

1:07:30

Craft, taste, and the AI "tell"

The host quotes Jorge Manrubia's "Oh my craft," which says the need to intervene on small details is fading as models improve and that he hasn't opened a code editor in months. Is craft decaying? Appleton says she doesn't care much about clean code, but she does care about design, and agents can't yet meet her standards. She is still deep in the details, correcting transitions. She has tried to write skills with her personal rules (padding between border and element, border shadows), knowing they aren't universal. The agents don't apply them consistently, so she still adjusts values by hand that she would have expected to be automated by now.

She isn't sure how she would feel if the skill suddenly worked. It might mean "all the fun bit is gone," since a gorgeous interface she didn't shape may not be satisfying. She compares it to Pinterest, where many interior design images are now AI-generated. They look beautiful, but they aren't real rooms, the light may be wrong, and so they're nearly useless for designing a real space.

She predicts that if agents can easily reproduce today's popular aesthetic (she names Linear and Vercel: minimal, clean, white, slightly rounded corners), it will become a sign that an agent made it, and people who want to stand out will do something new. Design has universal basics like readable text and enough space, but much of it is fashion. MySpace was once in style. Linear's look will seem dated in ten years. Models don't understand that design carries cultural meaning that changes over time. She says Claude has a recognizable style (cream backgrounds, reddish text, eyebrow labels), and people may reject clearly agent-made products out of hand.

The host says they can immediately spot AI-written text but wonders whether only experts notice. Appleton, also a writer, closes the tab at the first AI-sounding sentence. She describes reading a nonfiction book from 50 years ago and finding it refreshing to see none of the "it's not X, it's Y" patterns.

1:14:22

Digital gardens, home-cooked software, and barefoot developers

Her digital garden, started around 2020, is a blog where posts don't have to be finished as long as readers are told. Posts are labeled seedling, budding, or evergreen and show when they were last updated. Some are three paragraphs. Some long essays have a "draft in progress" marker, below which the text is rough notes. One essay she began three years ago moves forward a paragraph at a time. She has perfectionist tendencies and says she would publish almost nothing otherwise. She sees it as a contract with readers: rough work is fine if it's labeled.

At the 2024 Local-First Conference in Berlin, she built on Robin Sloan's idea of "home-cooked software," software made for yourself and your family like a home-cooked meal. His example was a video-messaging app for his family, not on anyone else's servers and not run by a company. She predicted an explosion of this with language models. She says it has now happened, with people building their own recipe managers, household apps and gym apps.

She extended it with "barefoot developers," after the barefoot doctors of Mao's China: rural villagers given basic medical training such as vaccinations and antibiotics, which she describes as a very successful program. Professional developers are expensive, but communities like an allotment garden or a street have small needs that Google Sheets doesn't cover, such as scheduling or inventory. Otherwise they pay for software that almost fits or trade away their data. A power user who once would have used Notion or Airtable could now build what they need with an agent cheaply, without reading code, as long as they verify it works. The host compares this to 1990s webmasters, often volunteers who ran school or company systems and helped each other in forums.

Appleton says such people already exist but need better support. Vibe coders often lack foundations and hit security problems, mishandle data, or lose databases. She wants strong local-first frameworks for them, with data on the device, sync between machines, no needless cloud database, and data ownership. These would include good security, reliable persistence, and good interface primitives, maybe as "the next version of Airtable."

1:21:06

Advice for engineers, and lessons from anthropology

Her advice to engineers mirrors her advice to designers. Just as a designer can ask an agent to explain the shape of a back end, an engineer can ask it about the principles of a good sidebar, typography, line height, or line length, or how to run user interviews and usability tests. That knowledge used to be hard to get outside a design team. She thinks agents are decent at this conceptual teaching, even if they don't always get the interface itself right.

From anthropology, she suggests looking at problems through their unspoken cultural rules: how users decide something is trustworthy or worth their time, or how a flow should go. She notes that some cultures see time as flowing right to left or top to bottom, and some group colors differently, for example treating blue and green as one. The point isn't always direct application, but knowing how widely people's interpretations vary, and that tools can change how people see the world. The host suggests engineers embed themselves with customers as anthropologists do. Appleton agrees. As engineers gain some free time, she says, they could move toward user research: understanding the domain, the context of use (on a factory floor, on the Tube), and when someone picks your product over another.

Her book recommendation is Addiction by Design by Natasha Dow Schüll. It is an anthropological study of gambling machine addicts in Las Vegas that also looks at the machine designers: how a machine keeps someone seated for 12 hours, and why casinos have no windows. She read it at university and says it makes you think about phones, Instagram, and "what are these systems we're building for people."

In closing, the host reflects that Appleton, even while prototyping the future of agentic tools, starts with pen and paper because "agents cannot see." They wonder whether that applies beyond visual work, recalling their best whiteboard sessions. They compare her uncertainty about agents producing perfect designs to their own feeling that writing code has become more transactional. They end by saying they might start keeping a notebook for ideas, architectures and systems, as some AI-free space.