Maggie Appleton on Design Engineering: Paper, Jigs, and What Agents Still Can't See
The Pragmatic EngineerMaggie Appleton is a staff research engineer at GitHub Next, where she prototypes ways software engineers might work with AI agents. She has been an illustrator, the first designer at the AI startup Elicit, and a front-end engineer. In this conversation on the Pragmatic Engineer Podcast, she explains what a design engineer is and why designers need to understand engineering constraints. She describes how her design process has changed with agents, and why she still starts nearly everything in a paper notebook. Her position is that agents have made implementation and prototyping much faster, but the early, visual, judgment-heavy part of design remains a human job. She is not sure how she would feel if that changed.
From cultural anthropology to front-end engineering
Appleton studied cultural anthropology at university and calls it her "one true love." She describes it as the study of human beings through participant observation: living closely with people to understand how they see the world. What drew her in was how much cultures differ, not just in food or housing, but in how they understand color or time. She grew up as an expat child who left London at six, and the field confirmed something she already felt: there is no fixed way to build a society, and the range of what is possible is wider than people assume.
When graduation came, the career options her professors described were a PhD and professorship, or working for the military. She went elsewhere. As a 90s kid she had learned HTML and CSS around age 12 or 13 through Neopets and MySpace. She had also done IT work for her university. So she started freelancing in web design to pay rent, and gravitated toward illustration because she loved drawing.
Her first exposure to product design came at a design and development shop in Prague that built apps for San Francisco startups, including early work for Tinder and Uber. She made illustrations, logos and branding there, and realized "there are people who design buttons and sidebars." She then spent about four years as an illustrator at Egghead, the developer education company, and later became its art director. She says that was her real front-end engineering education. To illustrate React components, useEffect, and JavaScript functions, she had to understand them, and she gradually moved into front-end engineering and visual design.
Elicit: a feature a week for a year
Around late 2021 or early 2022 she joined Elicit, an AI startup that uses language models to speed up scientific literature reviews. That work traditionally means academics reading thousands of papers and copying data into spreadsheets, which she calls a perfect use case for models. One founder had studied language models for about ten years and was convinced they would change science.
The team had six or seven people when she joined and stayed small. She was the only designer the whole time. There was a PM for a while, who left and was not replaced. Elicit had a popular free prototype, which gave the team a lot of data on searches, clicks, and A/B tests. She says they shipped a feature every week for a year, in a loop of design, build, ship, measure. Over a little more than two years she learned product design end to end, from user needs and domain understanding through front-end implementation. She describes the experience as wonderful and intense, and says she left wanting "to step back from startups for a minute."
What a designer actually does: nouns and verbs
Asked what a designer does, Appleton says it is not that different from engineering. It is problem solving with different materials. The process is the same: define and scope the problem, explore possible solutions and their trade-offs, prototype, and validate with users. The materials are space, size, weight, color, motion, prominence, and copy, for example whether a button's words tell the user what it will do.
In product design, she adds, you also work with the "nouns and verbs" of a product. For a sneaker store the nouns are easy: sneaker, cart, money. For something like AWS they get very hard. She has mostly worked on power-user tools for scientists and developers, where the nouns become abstract: the right container for a set of data, how one dataset connects to another, what to call a function that transforms data. Developer tools are especially hard because they are so malleable, meaning they could take many forms. A government form is a hard design problem but not complex in the same way.
Her current work on agentic tools shows the problem. Nobody yet knows the nouns of agents. There are chat sessions, plans, MCPs, skills. A designer could create a new primitive that connects several sessions to a pull request, and each choice of boundary changes how users think about the product. Then come the verbs. Editing, renaming, and deleting are standard CRUD, but can you fork an agent session, and what would that imply? For her, the hard part of product design is reducing all that possible complexity to a coherent, canonical set of nouns where users can point at something, know what it will do, and not be surprised by what happens when they click. She says this is "very easy to say" and extremely hard to do.
The host compares this to Kent Beck and Ward Cunningham keeping a thesaurus nearby while naming the constructs that became design patterns. Appleton agrees. In tech, she argues, product design is really software design, and software design is engineering. A designer who cares about colors, shadows and motion also has to understand the database shape, the APIs, and how data flows through components.
Who counts as a design engineer
The host pushes back with experience from Uber and Skype. There, designers worked in Sketch and later Figma, ran explorations and user testing, and handed visuals and animations to engineers without knowing how the database worked. Appleton says her view reflects her own background. Not every domain requires deep technical engagement. For government forms, a designer probably doesn't need to understand the database and is better off advocating for users through interviews and flow design. But for developer tools, or in a new field like AI where the back end largely decides what the product can do, designers have to engage with the technology. She sees a spectrum from business and market context through users and interface down to the back end. Designers can be valid anywhere on it. She likes the slice that is half engineering and half design.
A few years ago she tweeted about wanting to become "a full-blown design engineer" and collected names of people doing that work. She now thinks the role exists, but defines it more narrowly than much of X does. Many people who call themselves design engineers there are, in her view, very good micro-interaction designers: hover effects, loading transitions. She says that work is cool, but most of it can now be done by telling an agent without reading any code, so she sees it as advanced motion and visual design rather than design engineering, and doubts it is a full career.
For her, a design engineer is still a designer, caring about product nouns, verbs and visuals, who also works closely with engineers and usually implements much of the product directly, with a deep grasp of its technical architecture. The questions are things like: given the shape of the back-end data, what is possible in the interface? Given what the models can do and what custom skills the product has, how do you explain its capabilities to users? She has always done her own front-end work. Engineers she has worked with were usually thrilled, she says, because they hate CSS and could focus on sync and logic instead of whether a border radius is right. They rarely care about that, and she does.
Designers must understand their materials
Because she has always been the front-end person, Appleton says she hasn't felt the classic tension where a designer delivers a perfect Figma file and the engineer doesn't build it to spec. The host has. They recall designs with gradients that would have wrecked scroll performance on mobile, and designers focused on iOS while engineers worried about low-end Android devices.
Appleton blames the tools more than the designers. Sketch and Figma pixel mockups, she says, "have no relationship to the constraints of the medium." Any designer has to understand their materials. A table designer knows how oak behaves and how pine dents. A web designer who doesn't understand performance, data fetching, loading times, or race conditions will produce bad design solutions and bad working relationships. She hopes agents will help here. Designers can use them to move toward engineering and as tutors: "the engineers told me about this error state I've never heard of — explain why it would occur and help me build a state machine." She finds it hard to imagine things going well for designers who still want to make only pixel mockups.
Her toolkit: everything, plus paper
Her tools change constantly, partly because her job at GitHub Next is to figure out what GitHub should do next. She dogfoods internal tools and tries every agentic harness she can: Codex, Claude Code, Conductor, Amp, OpenCode, Pi. She currently loves Codex and calls OpenAI's desktop app beautiful and well thought through. She still uses paper and pen for early brainstorming, and Figma because she knows it well.
Once she has worked out what to build, which she says is most of the work, she hands it to an agent. Implementation isn't "solved," but it is good enough that if she specs something well and lists how the agent should verify its work, she can say "let me know when you've got a PR up." She rarely looks at code now beyond skimming PRs in review. She opens VS Code occasionally. She notes that GitHub Next builds prototypes to validate ideas, not production software, so its quality bar is much lower than github.com's.
Sketching a better way to plan with agents
Appleton showed her notebook on camera. Her current prototype starts from a belief that planning is a bottleneck with agents. Today planning usually means a long chat, often in a CLI, where the agent asks a series of multiple-choice questions, sometimes a hundred of them, for example with Matt Pocock's "grill me" skill. The host says they went through 36 questions and got annoyed around question 20 or 25, with no idea when it would stop. Appleton says people tire and can't make that many decisions that quickly. They also lack enough information for each one, so when an option is marked "recommended" they just start accepting it. Agents love to produce reams of text, which is a good output for agents but "not an ideal input for humans."
Her sketches explore a graphical interface where each decision gets its own small document or card. Simple decisions can stay multiple choice. Others need to show more. If the agent asks whether borders should be 10%, 12%, or 15% gray, it should show the options visually. If it asks how to structure the architecture, it should show an architecture diagram, data-flow diagram, or state machine. She is prototyping decision cards with embedded diagrams, prototypes and HTML, nested inside a larger plan document. The interface should be multiplayer so a team can plan together. Each decision should have a named human owner for an audit trail. The point is not to hunt people down, but so that six months later someone can open the card and see what information was available when the choice was made. On paper she is trying out shapes: an expanding stack of cards, one long linear document, or swiping like Tinder.
The notebook also has lighter projects, including a character for an ESP32. She credits Steve Ruiz with persuading everyone to buy these small Wi-Fi and Bluetooth devices with screens. She wants it to tell her the weather and how much agent capacity she has left before her usage limit resets.
Why pen and paper aren't going away
Her sketching habit comes from illustration, where you have to draw to work out composition. She trained in LA with film concept artists whose method she calls very technical: building things from 3D shapes and understanding how robot joints work so a drawn robot is believable. That background made UI sketching come naturally.
Her argument for paper is about speed and looseness. She could tell Claude Code, "here's an idea, it's like a stack of cards and it's an accordion," but drawing it by hand is faster and takes much less effort. Early on you need quick, loose feedback to find the shape of something before you can put into words what you want an agent to do. A sketch also stays on the desk, so the next day she remembers what she was working out. The sketches don't have to be visual designs; data diagrams work too. The key is that they externalize thoughts that aren't linguistic.
She says agents mostly take text as input. They can read images, and she does give them photos, but they are weak at spatial reasoning and visual design. They leave out spacing, size things wrongly, and overlap text: "They can't see." So she still does much early design without agents, and is skeptical of people on Twitter who claim to involve agents earlier.
The host compares this to engineers at a whiteboard, where drawing boxes and arrows makes people put their ideas into shared 2D space. They add that Miro caught on by bringing that to remote teams. Appleton connects this to how primitive agent interfaces still are, only about four years after ChatGPT launched in November 2022. She recalls that at Elicit they had been debating a chat interface when ChatGPT came out. Software design itself is a young field, she says, perhaps 60 years old. Agents live in a world of weights, skills and MCPs, and humans live in one of physicality, texture and light. The hard challenge is finding artifacts where the two can meet. She is clear that she doesn't think agents are conscious. She calls them a kind of intelligence that thinks and acts differently from humans. Her "eventual dream" is an agent that can look over her shoulder at her notebook and help move her ideas forward, but she thinks that is a long way off.
She has also started woodworking. Partly it's a joke that every software engineer eventually picks a physical hobby like ceramics, bread, or motorcycle repair. But she also bought an old London house that needs shelves and a repaired banister. She used Claude, ChatGPT, YouTube and Instagram as coaches, then signed up for a course, expecting parallels to software "in a more satisfying way where you actually touch the thing you make."
Jigs: building your own Figma
Appleton uses Figma to move past notebook sketches to medium fidelity: exact contrast, text sizes. She never takes it to high fidelity, because things look different in the browser (font rendering, responsiveness, breakpoints), so she moves to a browser prototype early. Now she points an agent at the Figma file to build it. For high-fidelity mockups she has run a loop overnight: take screenshots, compare with the Figma file, keep going until they match. By morning the interface is roughly as specced. She calls this grunt work ("we're building another sidebar"), where code quality and tech debt don't matter because the prototype may be thrown away.
Her favorite technique is the "jig." The term comes from woodworking, where a jig is a small device built for one specific job. In a live prototype she tells the agent which values she's unsure of, such as headline size, colors, or an animation curve. The agent adds sliders and color pickers so she can tune them live, then commit the final values. She calls it "build your own Figma as needed," not limited by Figma's features.
She demoed a prototype for a new GitHub Next website: an animated, spinning constellation of the team's past projects, with physics on the stars, hover scaling, and a scroll transition into the page. She began with a notebook sketch, then built jigs for background color, rotation speed, gravity intensity, star variation, resting logo size on scroll, and the strength of a glass effect. She showed how too much glass washes out the text and too little makes the effect invisible, so the right value is somewhere between. She says she wrote no code and didn't need to know the physics. She has a skill that tells the agent how to make these controls look decent. The one in the demo used DialKit. Static mockups of this in Figma didn't work, she says, because you can't feel the layout or scroll distance until it is live.
She links this to Bret Victor's 2010–2013 talks, including "Stop Drawing Dead Fish." His point was that programming isn't live: you edit variables in code, build, and look at the result elsewhere. He argued for direct, instant feedback on the artifact. In her view, jigs finally make that possible. Before, the closest thing was browser dev tools, and hand-building such controls wasn't practical because coding the thing itself was slow enough.
"Do we still need designers?"
The host quotes an engineer who asks Claude to generate 20 high-fidelity alternatives, iterates to a "perfect design," and finds it faster and more enjoyable than working with a design team. Appleton thinks that's fine when no designer is available and you need an interface to validate a hypothesis, especially for something simple like a button and a sidebar. Taste matters, though, and she might see those designs as obviously AI-generated and unlikely to do their job. When you are defining a new primitive, you need someone whose job is to think the problem through: run experiments, show them to users, and learn what people actually understand.
When models design for her, she judges the quality as poor. She thinks they've been prompted with universal design principles but lack nuance and context, so they apply the principles too rigidly. Her example is labels. Where a designer would use an icon button to close a sidebar, because users know the convention and will try clicking it, the agent writes "Close sidebar" or "Close modal," or adds four lines of instructions above a button. She understands the model is trying to make the interface explainable, "but actually it's really bad design."
The UI of AI may just be a table
Asked what good UX for AI looks like, Appleton says the question won't be answered for a while, but many principles haven't changed. At Elicit, scientists extracted paper data into Excel, one paper per row. The team thought that was old-fashioned and tried infinite canvases of scattered cards, linear card stacks, and Notion-like composable documents. In every interview users found them confusing and asked, "Can I just have a table?" After a couple of months the team concluded the new UI of AI was, at least for their case, a table. The familiar interface had the lowest cognitive load. They later built many features into and around the table.
She expects AI interfaces to keep using documents, sidebars, cards and tables, starting from familiar primitives and expanding outward. Chatbots weren't the final form. Codex still has a chat window but adds git worktrees, annotations, and debugging around it. The host adds that every new paradigm has a learning cost, like teaching parents or grandparents to use phones, and people don't want to pay it again.
Capability gaslighting
Appleton coined "capability gaslighting" before reading Ethan Mollick's similar "jagged frontier." Models are very good at some things and very bad at others, and it's hard to predict which a given task will be. Their impressive successes convince you they're highly capable, then they fail at something else. She says she sometimes misses how badly a model has failed because she still believes "Opus could never really get this wrong," and it "totally does all the time."
She contrasts this with human experts, who rarely lose their expertise suddenly. If they did, you'd wonder if they were having a breakdown. Models can do a task well one day and fail the next because of different prompting, different context, or plain stochastic variation. She grants that frontier models have a floor they won't drop below, but says upgrades can shift what a model is good at in ways that are hard to predict.
One developer, two dozen agents, zero alignment
Her talk with this title describes a problem GitHub Next has worked on for more than a year. Agents make each person much faster, but software is built by teams that need to agree with product managers, designers and other engineers. Slack, Linear, and GitHub Issues exist, but much planning happens before an issue is written, because by then you're usually ready to hand off to an agent. Teams need to agree on whether to build a feature, its shape, its interface, and any database migrations. She says there are few good tools for that stage, especially ones involving agents.
The biggest gap, she says, is that agentic coding tools aren't real-time multiplayer. Sessions are private and often local. That is starting to change. She mentions Buzz from Jack Dorsey and GitHub Next's own ACE prototype, which offered shared compute and sandboxes in a Slack-like interface so people could code and talk at once. But even with an agent in a chat channel, teams still need ways to agree, which is why she wants decisions to become a first-class primitive: large enough on screen to inform the choice, open to teammates' input, owned by someone, and recorded as context for the agent. Before, implementation took long enough that teams could adjust along the way. Now there is a hard handover to an agent, so alignment has to happen up front, and she says her own team struggles with this too.
ACE combined Slack-style chat with cloud microVM sandboxes, PR creation, and code review. She says it was a useful prototype but too ambitious for a team of three or four, and couldn't become a full product. Parts of it are being shipped. Sandboxes are now part of the GitHub desktop app. She thinks it also made leadership take multiplayer more seriously. The host describes Ramp's internal agent, Inspect. It works in Slack, a Chrome extension, and a web interface, and all sessions are public so anyone can join and redirect them. Ramp worried about privacy but found it wasn't a problem, and the openness helped adoption. The host credits its success to deep integration with Ramp's internal systems and full cloud developer machines, which is easier for one company than as a generic product.
GitHub Next's role, she explains, is to run ahead of the product teams and explore bigger or riskier bets. After multiplayer sessions, her team is now looking at proactive agents that aren't annoying. They start from an assumption that intelligence is nearly free. If you had 100 background agents, what is the least overwhelming way to consume their output? How could agents work in a shared document without feeling invasive? She calls these open research questions and hopes the prototypes will inform GitHub if it builds in this direction.
Craft, taste, and the AI "tell"
The host quotes Jorge Manrubia's "Oh my craft," which says the need to intervene on small details is fading as models improve and that he hasn't opened a code editor in months. Is craft decaying? Appleton says she doesn't care much about clean code, but she does care about design, and agents can't yet meet her standards. She is still deep in the details, correcting transitions. She has tried to write skills with her personal rules (padding between border and element, border shadows), knowing they aren't universal. The agents don't apply them consistently, so she still adjusts values by hand that she would have expected to be automated by now.
She isn't sure how she would feel if the skill suddenly worked. It might mean "all the fun bit is gone," since a gorgeous interface she didn't shape may not be satisfying. She compares it to Pinterest, where many interior design images are now AI-generated. They look beautiful, but they aren't real rooms, the light may be wrong, and so they're nearly useless for designing a real space.
She predicts that if agents can easily reproduce today's popular aesthetic (she names Linear and Vercel: minimal, clean, white, slightly rounded corners), it will become a sign that an agent made it, and people who want to stand out will do something new. Design has universal basics like readable text and enough space, but much of it is fashion. MySpace was once in style. Linear's look will seem dated in ten years. Models don't understand that design carries cultural meaning that changes over time. She says Claude has a recognizable style (cream backgrounds, reddish text, eyebrow labels), and people may reject clearly agent-made products out of hand.
The host says they can immediately spot AI-written text but wonders whether only experts notice. Appleton, also a writer, closes the tab at the first AI-sounding sentence. She describes reading a nonfiction book from 50 years ago and finding it refreshing to see none of the "it's not X, it's Y" patterns.
Digital gardens, home-cooked software, and barefoot developers
Her digital garden, started around 2020, is a blog where posts don't have to be finished as long as readers are told. Posts are labeled seedling, budding, or evergreen and show when they were last updated. Some are three paragraphs. Some long essays have a "draft in progress" marker, below which the text is rough notes. One essay she began three years ago moves forward a paragraph at a time. She has perfectionist tendencies and says she would publish almost nothing otherwise. She sees it as a contract with readers: rough work is fine if it's labeled.
At the 2024 Local-First Conference in Berlin, she built on Robin Sloan's idea of "home-cooked software," software made for yourself and your family like a home-cooked meal. His example was a video-messaging app for his family, not on anyone else's servers and not run by a company. She predicted an explosion of this with language models. She says it has now happened, with people building their own recipe managers, household apps and gym apps.
She extended it with "barefoot developers," after the barefoot doctors of Mao's China: rural villagers given basic medical training such as vaccinations and antibiotics, which she describes as a very successful program. Professional developers are expensive, but communities like an allotment garden or a street have small needs that Google Sheets doesn't cover, such as scheduling or inventory. Otherwise they pay for software that almost fits or trade away their data. A power user who once would have used Notion or Airtable could now build what they need with an agent cheaply, without reading code, as long as they verify it works. The host compares this to 1990s webmasters, often volunteers who ran school or company systems and helped each other in forums.
Appleton says such people already exist but need better support. Vibe coders often lack foundations and hit security problems, mishandle data, or lose databases. She wants strong local-first frameworks for them, with data on the device, sync between machines, no needless cloud database, and data ownership. These would include good security, reliable persistence, and good interface primitives, maybe as "the next version of Airtable."
Advice for engineers, and lessons from anthropology
Her advice to engineers mirrors her advice to designers. Just as a designer can ask an agent to explain the shape of a back end, an engineer can ask it about the principles of a good sidebar, typography, line height, or line length, or how to run user interviews and usability tests. That knowledge used to be hard to get outside a design team. She thinks agents are decent at this conceptual teaching, even if they don't always get the interface itself right.
From anthropology, she suggests looking at problems through their unspoken cultural rules: how users decide something is trustworthy or worth their time, or how a flow should go. She notes that some cultures see time as flowing right to left or top to bottom, and some group colors differently, for example treating blue and green as one. The point isn't always direct application, but knowing how widely people's interpretations vary, and that tools can change how people see the world. The host suggests engineers embed themselves with customers as anthropologists do. Appleton agrees. As engineers gain some free time, she says, they could move toward user research: understanding the domain, the context of use (on a factory floor, on the Tube), and when someone picks your product over another.
Her book recommendation is Addiction by Design by Natasha Dow Schüll. It is an anthropological study of gambling machine addicts in Las Vegas that also looks at the machine designers: how a machine keeps someone seated for 12 hours, and why casinos have no windows. She read it at university and says it makes you think about phones, Instagram, and "what are these systems we're building for people."
In closing, the host reflects that Appleton, even while prototyping the future of agentic tools, starts with pen and paper because "agents cannot see." They wonder whether that applies beyond visual work, recalling their best whiteboard sessions. They compare her uncertainty about agents producing perfect designs to their own feeling that writing code has become more transactional. They end by saying they might start keeping a notebook for ideas, architectures and systems, as some AI-free space.
Like you have an idea in your head and sure I could go into Claude Code and be like, "Hey Claude, here's an idea I have." It's like a stack of cards and it's an accordion, but it's much faster to just get a piece of pen or pencil on the desk next to me and draw that with my hands. And then you can look at it. It doesn't go away on your screen. It can sit on your desk and you can like the next day be like, "Oh yes, I remember."
Like you have to understand the materials you're building with in any design role. If you're designing for the web and you don't understand performance or like how is your app fetching data and what's the loading time and what if there's race conditions, you'll end up with bad design solutions.
I'm hoping agents actually help solve this because then designers using agents can step more into the engineering side. So, you can play around with like layouts and ideas, but you can't get a sense until it's live in the browser of how this is going to feel and like how much scroll you need to do here. Whenever I'm designing something visual, it's so much easier to have live variables, which again would not have been possible before.
You did a talk titled One Developer, Two Dozen Agents, Zero Alignment. Why we need collaborative engineering?
Us alone with the agent, we can go really fast. The software is always built on a team. The biggest gap seems to be like we have agentic coding tools, but
What makes for a great designer or a great design engineer and what can us software engineers learn from them? Maggie Appleton is one of the most thoughtful design engineers I know. She's worked as an illustrator, as a designer at an AI startup, and is currently prototyping at GitHub Next.
Today we talk about what is a design engineer, and why there's often a tension between engineering and design, the tools she's used before and now, and why pen and pencil is not going away, even with agents. Do you still need a designer when AI can generate 20 high fidelity prototypes, and is AI design obvious to spot? And many more. If you're an engineer wanting to know where design is important and how to get better at design yourself, this episode is for you.
This episode is presented by turbopuffer, vector and full-text search built on object storage. It's fast, cheap, and extremely scalable.
Today's episode will be about design and design engineering. And so I wanted to share something visually interesting about our season sponsor, Antithesis. We already know that Antithesis verifies your system correctness by running your whole system in hostile simulation and finding bugs.
Here's a UI for casualty analysis. You can open a report for a bug and see the probability of a bug occurring throughout the timeline of the simulation. In this case, we can see that at virtual time 25, something happened that makes this bug close to 100% to occur. So, we can jump to this point in virtual time simulation to read the logs. This kind of bug probability visualization is one that I've never seen before.
There's also this neat log explorer. You can filter on error messages and then visualize how common or uncommon the error is over time. For example, here's looking for failing linearization failures, the purple line, and you can understand how rare or common a specific failure was. Again, I've yet to see this kind of error visualization, and I really like the innovation on the UI side.
And finally, the simplified multiverse debugger. You can go back in time and replay a debug timeline. And you can inject bash commands at any time without affecting the playback of the bug. How cool is that? For example, we're listing files in the current directory, but as you can imagine, you can go debug the whole environment so much easier. I love how the team at Antithesis are pushing what's possible with both debugging and verifying software. Head to antithesis.com/pragmatic to learn more.
Maggie, welcome to the podcast.
Thanks. I'm thrilled to be here.
I've had so many guests on the podcast. No one had quite the introductory story into tech like you. How did you get into tech? And you came from a very different background, right?
Yes. Yeah. I mean, I came in because it was where the money was, I guess. I mean, not really, but in the sense that like in university, I studied cultural anthropology, which is like my one true love. Like I adore cultural anthropology. I totally fell in love with it in university.
What is anthropology?
Yeah, okay. It's the study of human beings, which sounds impossibly broad and like how could that be a discipline, but it does it in a particular way where you go and you live with people intensely in this thing called participant observation. Usually it was done kind of when the field was born of different cultures. Of course, it was like anthropologists from the west primarily going to places like Papua New Guinea or Australia and living among kind of traditional peoples and then realizing how different their cultures were. Not just like, oh, they eat different food, you know, they have different houses, but they have completely different understandings of like what color is. They have completely different understandings of time. It was kind of alongside the birth of psychology like understanding how flexible is the human mind about constructing understandings of the world.
And anthropology is really one of these eye opening subjects when you get into it because you can kind of get into like medical anthropology or the anthropology of sex and gender and you just find out how extremely adaptable and fluid human beings are.
And I loved it because I grew up as an expat kid overseas. So I think I was exposed early to lots of different cultures. And so it felt very natural to me to realize like, oh, the way people do things at home, you know, quote-unquote home would be London for me, but I left at age six, is completely different to the way they do it elsewhere. And there's no fixed way for humans to kind of construct a society or live life and the bounds of what we think is possible is much wider than we originally assume, which is what I loved about it.
So I studied it, but of course come senior year kind of go, right, what jobs are there available in cultural anthropology?
I guess living with natives isn't all that many.
It's not very lucrative. So the options were like go get a PhD and become a professor or the military hires lots of anthropologists to come up with torture techniques for people in different countries. So when presented with these options by our professors, we were like, "Okay, okay, we'll go think about that."
And I had always loved design growing up. So I had been doing like, you know, I was a kid of the 90s. I grew up on Neopets with HTML and CSS and MySpace and I learned HTML and CSS probably around 12 or 13 and knew how to do it but in the way where like there wasn't that much complexity to it in whatever this was, 1999.
Did you have a MySpace?
I did. Oh. Oh yeah.
Oh. And you customized it with all the
Crazy animations and like the little sparkle trail at the end of your cursor and like I think mine was quite goth at the time. But I learned a lot. But of course, at the time, the web was so new. It wasn't like people thought this was a career or you didn't really think of it that way. But I came out of university with this degree that wasn't necessarily employable. And I just started doing freelance web design work because that was how I knew how to make money. Like throughout school, I was doing IT tech work for the university and just naturally went into this because it was like, well, I need to pay rent somehow.
And I gravitated towards illustration originally cuz I love drawing. So I was originally an illustrator for the first couple years.
And the way that I went more into the tech side, because you can kind of be an illustrator that's like, you know, editorial or you could go into branding or like all kinds of types. But I started working for a startup. I guess I wouldn't call it a startup. It was more like a design dev shop in Prague that was making UI/UX apps for like SF startups at the time. They were doing work for like Tinder and Uber like in the early days, like in their beginnings. So it was there I got exposed to UI/UX design and I was making the illustrations for these apps and doing logos and branding. But that was my first exposure to like, oh, there are people who design buttons and sidebars and that's kind of interesting. I didn't end up going into that for a while but that was my first introduction to like what is product design in tech as a field.
And eventually I joined this company called Egghead which does developer education. So they taught JavaScript. Sure you've seen them around. And I was their illustrator for four years. And then I became an art director and like art directed other people's illustrations for them. And that was really where I learned JavaScript, React. Like that was I think my real front end engineering education, was like doing illustrations for them. But in doing the illustrations, I had to understand the material I was illustrating, which turned out to be like React components and useEffect and like JavaScript functions and in a weird way, I just ended up moving more into front-end engineering and visual design because it just felt like a natural move.
You worked at Elicit as well, right? Was that before Egghead?
After.
That was after. Yeah. So you were at Egghead as their learn design there and then Elicit was an AI startup, right?
Yeah. Early. Pretty early. Well, okay.
Pretty early. Yeah.
I mean, it was like the founders there were really kind of incredible people and one of them had been studying language models for 10 years. He had like seen the writing on the wall way before anyone else, but like kind of out of MIT, like PhD kind of like, you know, machine learning stuff and kind of had this realization of like language models are going to revolutionize science. He was really big on like how could this speed up the scientific process.
So Elicit originally was, and it still is actually, I mean they've expanded, but at its core it is using language models to speed up the scientific literature review process, which is usually a very manual slow thing of all these academics reading thousands of papers and extracting data about them into spreadsheets. Perfect for models, like a perfect use case. So I joined them. I think it was late 2021, early 2022.
And you were the only designer. There's no product manager. You were the founding designer.
Yeah. Yeah. I think there were six of us in the beginning when I joined, maybe seven. I might have been the seventh. I forget. A very small team. And it stayed small most of the time I was there. And yeah, I was the only designer the whole time. And then we had a PM at some point who left and then we didn't get another one. So there's no PM for quite a while. And it was a really wonderful experience cuz it was a classic early startup. Like it was like we were family. We would like stay together for weeks at a time during retreats. Like the founders had really strong conviction and they're really smart people and I trusted them so much. So it was sort of like get on board with the vision kind of deal.
And it was there I really learned product design end to end, I'd say, in the terms of we had lots of users cuz we had a free prototype that was like very very popular. So you had all this data you could collect about what people were searching for and clicking on and what worked and what didn't and A/B tests and we just went really fast. Like we shipped a feature every week for a year. So it was like design, build, like ship it, measure, design, build, ship it, measure on repeat over and over for like the solid, I think I was there a little over two years. And I had learned a ton in that time just about the core mechanics of product design in the sense of like really going from what are the user needs, what's their domain through to doing all the front-end engineering and like tying it all together. So it was a really wonderful experience. I learned a lot. It was intense. By the end of it, I was like step back from startups for a minute.
So you got into design I guess kind of more like self-taught, like figuring out there's a need for this. You worked at Egghead where you learned to design or explain developer concepts, educational concepts for developers. You worked at an AI startup as a designer. What does a designer do? It's kind of a simple answer now that you've actually several roles and I get a sense that of course like there's going to be the it depends, but at startups that you've worked with, that you know, how would you describe? And of course many of us developers have worked with designers, some have not at all.
I mean I kind of describe it, it's not that different to engineering. It's problem solving but the materials are different, is the way I would describe it. It's like you go through the same process of like defining what your problem is. Like are you sure this is the correct problem? Have you defined it well and scoped it well? You know, researching possible solutions, doing wide exploration. What were all the ways we could solve this? What are the trade-offs of them? You know, prototyping solutions, validating those are the right solutions. Do they work for users? Do they make sense? I don't think it's that different to engineering in the sense that I do some of both in my job, but it's just the material is different. Instead of working with code, although nowadays with developers you're working in higher level kind of architecture, data flow, you're not necessarily writing the syntax. But at the time, what syntax are you going to use to write this problem?
And with design, the materials are like space, size, weight, color, you know, things you'd see in an interface in terms of the visuals, motion, you know, prominence. Is this big enough? Are these the right words for the user to understand what this button's going to do? Like everything from copywriting to visual graphic design.
But with product design, you're also dealing with what we would call nouns and verbs of a product. So, it's easy when your product is like a sneaker store. It's like the nouns are like sneaker, cart, money.
If you're designing AWS, the nouns get extremely difficult. And I've primarily worked in what I'd call power user tools like scientists, developers, where the nouns are extremely hard because they get very abstract. It's sort of like what's the right container for a set of data or what's the right container or noun to point you to like, oh, this is your whatever set of data, or like this set of data connects to that set of data, or like here's a function that transforms this data into another set. You need a noun and verbs to give to users so they can understand how to manipulate whatever you're trying to get them to do. It's really difficult in dev tools sometimes because there's so much malleability in a way there isn't with stuff.
What is malleability?
Like ability, like it could take many different forms and shapes, versus if you're designing things for the real world, like I have friends who work for like government design. There's restrictions there where it's like you're trying to get someone to fill out a form. It can be a hard design challenge, but it's not complex in the same way. Versus like at the moment, right? I'm of course trying to design like agentic tools and it's like what are the nouns of agents? Like we don't know. There's like chat sessions. We have these things called plans. There's something called an MCP. There's skills. Like could we make new primitives that like connect a bunch of sessions all the way to a PR that becomes a new noun that contains it?
Like there's all kinds of boundaries you could draw that would make the user think about your experience differently. And then you have to define what verbs can they take on which nouns, right? Can you sure
edit, rename, delete? This is like classic CRUD stuff you have to figure out. But I don't know, can you fork an agent session? What are the implications of that? I think the hard bit of product design is designing a coherent system that takes all this complexity that could exist, especially in something like dev tools, and reducing it to the simplest possible form you can, which is very easy to say, extremely hard to do every time, to a really canonical set of nouns that the user can go, okay, I can point at that, I understand what that's going to do, when I click a button I don't get surprised at the outcome. This takes time, to do this kind of hard reduce it to its best, most elegant form.
This is so interesting because what you've described of, you know, we have a new product that has not existed before, like a dev tool in a digital space, which lives in our head or inside of a computer, with things that we just invented, may that be an MCP or an agentic skill, and then thinking of how we find the right words so people can use it and it makes sense. I just see that you're kind of drawing up a map in your head, which is not all that different to when you're building a new system as a software engineer. And I remember talking with Kent Beck, who talked about how with Ward Cunningham they came up with some of the basics of, that might have been domain-driven design or it might have been some related concepts, but they had a thesaurus in front of them, they searched for the right words on how to... oh, it was design patterns.
Design patterns later came out of it, but they were trying to put a name on these constructs and these things, which feels like a very similar thing to what you're describing.
Yeah. That's design. This is where I kind of get into, sure, there's types of design. There's interior design, there's brand design, there's product design, but product design when we talk about it in tech is actually software design, and software design is engineering. It's actually the same skill. You might be working with slightly different materials in the sense of one of you cares more about the colors and the size and the shadows and the shape of things and the motion design, but that person also has to understand what's the shape of the database, what are these APIs like, is the data in the right place, are we passing data through this component in the right way. They have to also understand.
This is interesting because the designers I worked with years back at places like Uber and at Skype, you know, they had their design tools, which was Sketch, later Figma. They typically worked with the product folks. They did explorations or UX prototypes. They often sat in on user testing, but in the end they had a design. They had visuals. They had animations. That was their thing. And they handed it over to us engineers together with the PR, or here's the product, here's how it's going to look. And then we built it. And we would build the UI, let's say, on mobile. And we would then go maybe sit with them or show them and they would say, "Oh, this motion doesn't feel good." But my view of a designer is, well, it was very visual, and for example those designers, maybe they didn't need to, but they didn't get involved in how the database works. They were very much aware of the flows, of the user journey.
So is that type of design, how would you characterize it as being different, or is it just that their domain is a little bit different, that's more at the business level, the more kind of mobile? These were mobile and web and some of those.
Yeah, I think it gets into how we label different parts of design. The kind of design I'm talking about is obviously biased by my experience, which is much more design engineering stuff. When I hear design engineer, that's kind of a catchphrase now, and what does that mean? But I think of it as a designer who really engages in the engineering and understands how the product works, and it's required. It's not required in all domains. Again, if you're doing government forms, I don't think you need to understand the database. But if you're designing for developers or you're designing in a brand new field like AI, where so much of how the product works is determined by what is possible on the back end and the shape of the data, then you do need to engage with it a lot. But if you're working on an app where actually a designer is probably better served by advocating for the user, I think that's a more traditional philosophy of product design: you represent the user and you represent trying to get the best experience possible for the user. And that means user interviews, caring about flows, like does this button feel big enough? Does it have the right words? And then they can spend much more energy there if they don't have to care about the technical back end. Maybe it's actually irrelevant, when really what you need to care about is does this flow make sense to people. So that's not necessarily a different kind of designer, but you kind of think of the whole stack, right? But expand it all the way out, not just back end and front end, but the interface and users, and then the product in the context of a business, and a product in the context of an economy. There are people who lean way more on this side. And I'm a bit more straddling the product design and the engineering. But you can have valid designers all along this, and some people, again, just visuals or just animation. There's lots of niches. I just like a bit more of the half engineering, half design slice of it.
To help understand a bit more of what you do, can you talk about the tools that you use? As engineers we're kind of used to our tools; it used to be the IDE and the code editor, and some of the hardcore people like Vim and some of those things, and these days of course it's changing a bit more, but those are the tools used. What are your tools?
It changes all the time. I mean, to some degree some things stay constant, and you're doing similar tasks, but of course at the moment I'm trying everything, because part of my job is trying to figure out what GitHub should do next. That's what our team does.
That's actually the name, right?
GitHub Next. Hey, it was well named. So a lot of the time, of course, I'm trying out Codex and Claude and dogfooding stuff internally at GitHub. I've tried Conductor, Amp. I mean, name any agentic harness, like OpenCode, Pi, I try them all. I do love Codex at the moment. I do think OpenAI is on to some really good stuff, at least in their design of their desktop app; it is beautiful. They've really thought it through. So I think they're doing some really good stuff. I still use paper and pen for initial brainstorming.
Yeah, you have your notebooks here.
Yeah, I brought some along because notebooks aren't dead. I don't think they're going to be dead forever. And I still use Figma because I know how to use it and I can do some brainstorming in it before I pass it off to an agent. But to be honest, there's a point where once you've figured out what you need to build, which is actually all the work, once you hand it off to an agent, it's not that implementation is solved, but we've reached a point where implementation is good enough that if I spec it out really well and I list out how the agent should verify for me that it actually did the work, I can hand it to an agent and just be like, "Right, let me know when you've got a PR up." You know, I don't really look at code that much anymore. I do look at PR code when I'm reviewing it, but skim, skim, skim, okay, that looks sensible. Merge. So most of the work I find is everything leading up to telling the agent to implement something, and then the PR review is kind of a separate piece of work. But the bulk of my tools deal in this first section now, versus we used to have the IDE, you know, once it was like, okay, you've decided this is the right thing to design, now you actually start the work of implementing it. It's kind of beautiful that that's no longer part of it as much. I open VS Code sometimes, but not really.
Yeah. But beforehand you would prototype, right? Like you build prototypes.
Yeah. As part of the what should we build, I include in this lots of prototyping, and I kind of have the privilege of being on a team where we don't have to build production-quality software. We are mostly trying to prototype and validate ideas. So we have a lower quality bar than someone shipping to proper github.com. That's a very high quality bar.
Yeah, which makes sense. You get the directional things right, and then once it's great, you might build it or you might decide to build it.
Yeah. But I definitely, I mean, we can hopefully show these later, I do a lot of stuff sketching. Even just interface design, a lot of what you're doing is just drawing boxes and then being like, does this slide over from the bottom if I click this button?
So this is you kind of drawing up how an interface you think could look. So I see a mix of UIs, and we'll put these onto the screen, but UIs, description of what they do. Can you just talk to one of these, like one that's interesting or memorable?
Yeah. So this is a new, I should explain the context. I'm prototyping at the moment something where my theory is that one of the bottlenecks with agents is planning; it's a really bad experience at the moment. At the moment you have a long chat with an agent, sometimes even in a CLI, which is a pretty primitive interface, and then the agent grills you by asking you a set of choice A, B or C questions, and it does this like a hundred times over if you're using Matt Pocock's grill-me skill or whatever you like.
He was just on a podcast, and yes.
And by question 20 you're quite tired and your brain starts shutting down.
That was my experience. I went through and I got 36 questions, and they were good, but I was starting to get annoyed around like 20 or 25. I'm like—
And also—
And I had no idea when it would stop.
Right, it's endless, and the human brain, you get tired, you can't make this many decisions in this short of time. And also you don't have enough information about most of those decisions, because it's given you a question and three options and it's told you A is recommended, and then you just start being like, yep, A, A, enter, A, I agree with you. So this is not an ideal experience. This gets into how agents love to output reams and reams of text, and that's an ideal output for agents, but it is not an ideal input for humans. So we have a mismatch with what humans need to be able to digest large amounts of information and truly understand it and be able to make informed decisions; that is not the interface for this. So I'm trying to explore how we would make an interface that got us to do better planning and better decision-making, but in a way that was easier for us to comprehend. So this is what I'm prototyping at the moment. And part of this is, okay, it's going to be multiplayer, because of course everything should be with your team planning together.
Yeah. So several people can do it together.
Exactly. But also, how do you make it so that you have more space for each decision? So I'm trying to prototype, okay, so first of all, this has to definitely happen in a GUI, not in a CLI.
So these are just rough ideas of how you could have more space, or—
Exactly. So I'm thinking of it as each decision maybe has its own little mini document or card. And depending on the decision, some of them maybe you can just have three multiple choices. That's fine. It's maybe a simple decision, right? A, B or C. But sometimes it'll ask me things like, do you think the borders should be gray 10%, 12% or 15%? And I'm like, well, show me; this is a visual question. Or sometimes it's like, how do you want the architecture to be structured? And I'm like, well, show me an architecture diagram. Show me a data flow diagram. Show me a state machine. I'm trying to prototype decisions that come with diagrams and prototypes and HTML embedded in them, so that depending on what question is being asked, it shows me the correct interface to make that decision. So what I'm prototyping here is, okay, you've got a plan, that's a big document, you've maybe got these little decision docs embedded within them that you can expand to see more of. And then I think each decision should have a human assigned to it who made that decision, so that you have an audit trail later.
Not necessarily to hunt people down, but to be like, okay, why did we make this decision about the back end? Well, let's go see. Okay, Luke made that decision 6 months ago, let's go open up his decision card and see what information he had available in order to make this decision. Hopefully this becomes useful later. But here I'm just sketching what are the possible shapes this could take. Is this a stack of cards that expands on the screen? Is this all one big linear thing? Are we swiping through stuff like it's Tinder? What interfaces are going to be useful for this? So a lot of the sketching is just trying to figure out possible shapes of things. I have all kinds. Oh, these are characters, because I was trying to make an ESP character.
Can you show the character?
Well, this is, oh, this is my kid scribbling all over my notebook. This is, you know, ESP32s. Do you know Steve Ruiz has sold everyone on buying an ESP32? It's this little device with a screen and a very small microchip that has Wi-Fi and Bluetooth. And so I was designing this little character who tells me the weather, and I'm trying to hook it up to my agent so it can be like, you have this much capacity left before your reset limit happens. So all kinds of things happen; it's a mix of some serious stuff and some fun stuff.
Yeah. And I can show the whole thing with notebooks and thinking on paper. I think all designers do this, but coming from illustration, the reason I do this is because I started in a world where you have to draw everything to figure out what you're going to do. So when I was an illustrator, I had tons of these notebooks where you're figuring out what's the composition, what's the physical shape of things, how are you working out the shape of the grass in a scene. And so I think I started problem solving on paper very early on in my career, figuring out compositions and layouts.
Was this in college? When was this?
I think I was working for Egghead when I was doing this.
And you drew all of these?
Yeah. I was in LA for a while, which is a terrible place to live, but I trained with people who are concept artists on films, and they have a really beautiful way of working that's very technical. It's very much constructing things from 3D shapes and drawing in space.
Because these are really 3D.
Yeah. And they kind of just teach you how to do landscapes and layouts. I loved their way of teaching, and again it's very technical. It's very engineering; you have to understand, like with robotics, different types of joints you could put together a robot in, so that you could actually draw a robot that was
believable. There was a lot of understanding reality in order to believably draw reality. So I think doing this set me up, when I moved into UI design, to do a lot of this kind of sketching because
It just comes naturally. I mean, these are a little bit easier to sketch as well in terms of the user interfaces.
Yeah. But then I think it helps you think on paper. All of this is just sort of like you have an idea in your head, and sure, I could go into Claude Code and be like, "Hey Claude, here's an idea I have. It's like a stack of cards and it's an accordion." But it's much faster to just get a pen or pencil on the desk next to me and draw that with my hands. It's way less effort. So I feel like people online keep being like, "Oh, everyone's just going to prompt everything to create." But you need something that is quick feedback and very loose in the early stages to figure out the shape of something before you can put into words what you want an agent to do.
And also it's visual. It's not text.
And is it not a bit more satisfying doing it with your hand?
So much better. Yeah. And then you can look at it. It doesn't go away on your screen. It can sit on your desk and you can, the next day, be like, "Oh yes, I remember I was trying to figure out what shape this feature should be." Or you can draw data diagrams, whatever you want. It doesn't have to be visual, but it's just a way of externalizing thoughts before they're linguistic.
I think this is a key thing: again, agents only accept text as inputs. Okay, they can read images. I do take in photos and put them in, but they're not as good at images. They're very bad at spatial reasoning. They're very bad at visual design. I mean, trying to get them to do design, they just make mistakes where they just don't put spacing around things and things are the wrong size and they make text overlap. They can't see, right? Trying to explain a visual idea in text to an agent is really challenging, and it doesn't work very well. So I find I still end up doing a lot of my design without agents up front, because all it is is thinking through the visual pieces of it. And then when I'm like, okay, I know this is the shape of the thing I want, then I can tell an agent to do it. But I find it hard to involve it earlier in a way that I think some people on Twitter are claiming they do, but I'm skeptical.
But this is interesting because even in software design, and I'm talking about architecture design, some of the, I guess, most productive sessions I've observed have been in person around the whiteboard. Yeah. Where, again, we're talking about components, databases, networking connections, retry logic, whatever. These things you could describe, or you could put it in a computer, but when you put it on a board, when someone puts it on there and then someone else takes it, they're forced to take their ideas into a 2D space, because we don't do 3D. I know you can draw cool 3D, but we cannot. So just boxes and arrows. Someone does that and the other people understand, and they go in and they add their own thing, or they circle, or they add new components. And even in digital space, I think Miro was a good example. They became so popular because they figured out a way to do collaborative whiteboarding. You don't have to be in the same room, but they give you somewhat similar tools.
Yeah. So I wonder if this whole thing of getting your ideas to a physical medium, which I think for us engineers is the whiteboard, for you it's the notebook, maybe it just helps to rethink or solidify. It also does have a forcing function about the verbs, the shapes that you
Yeah. Yeah. Yeah. And this gets into the interfaces we have to agents right now being so primitive. I think we know this, right? We're all a couple years into this entire thing, which is wild. Maybe 5 years. I forget when GPT-4 came out. Was it three or four years ago?
ChatGPT 3.5 came out... three years. Four years. Four years, November 2022.
Yeah. Yeah. Yeah. This was when I was at Elicit, because we were debating doing a chat interface and then they did it and we were like, oh, they stole our idea.
is necessarily competing with them. But that is no time at all, and it's kind of wild we've made it as far as we have. But I always say software design is an extremely young field. What are we, 60 years into it at most? So even there we haven't figured out a lot of things about how to design the best interfaces for people and machines to communicate with each other, and agents barely at all.
I feel like there's this world that agents live in. There's weights and models and skills and MCPs. And then you have your human side that is physicality and texture and light and materials and all these things agents don't understand. And trying to find artifacts that allow us to meet in the middle and create stuff together is the really hard challenge, because you've got two totally different types of beings. Not that agents are conscious beings. I'm not in that camp. But they're a type of intelligence that wants to think and act in a certain way that is not the way humans want to think and act. And it's hard to translate between the two. I just find myself very frustrated that they can't sort of, you know, be looking over my shoulder, looking at my notebook, and be understanding what I'm drawing and helping me move my ideas along. This is the eventual dream: they understand space and light and shape and lines. But I think we're quite a ways away from that.
So, we talked about your notebooks and digital tools, but one thing that you've posted about recently is woodworking. You said you're at that stage of software design where you start taking woodworking courses. Can you tell me a little bit about that experience?
Well, the little bit of the joke is I just feel at some point every engineer has a choice of things you can get into, because you feel like you're disembodied from the world by working in software. So you have to pick ceramics, bread baking, you can maybe pick motorcycle repair or something, but woodworking is a pretty popular one because it's like engineering. There's a lot of measurement and being precise. So yeah, I'm learning woodworking because I bought a house here, and all the houses in London are very old. So you buy a house that needs a lot of work. That is always what happens. Maybe the same in Amsterdam, I think. Old housing stock.
So then you suddenly become like, oh, I need to learn DIY skills, which, you know, Claude and ChatGPT are very helpful coaches in this regard, along with YouTube and Instagram. You can learn a lot. But I just was like, oh, there's so many things I need solutions to that I need to figure out how to make. I need to make shelves here. I need to fix this banister. So I was like, well, I have no DIY skills, no woodworking, but I just signed up for a course, being like, well, I'll learn. I'm pretty sure I'm not too old. I can acquire new skills, and I'm sure I'll find lots of parallels to software, but in a more satisfying way where you actually touch the thing you make.
One thing that you've been talking about is the concept of design engineers, from a few years ago. You posted, I'll quote you: "I'm having a strong 'should I just become a full-blown design engineer' day. I'm not even sure what it means, but I just want to touch lots of code and solve tangible problems on screens and make beautiful animated stuff."
Yeah.
And then you went over and you tried to collect names of people who you knew who you felt were kind of doing this design engineering. Now that was a few years ago.
Yeah.
What have you figured out about design engineers, if they exist? What they do? What places they work in.
Yeah, I think they do exist. I think Twitter might have a different definition, or I think there's a lot of people where... not that I should call it X. Sorry, X.
Twitter, X, all the same.
On X, I feel like a lot of people who get called design engineers, or who present themselves as design engineers, actually are very good micro-interaction designers. Little things like, here's a cool hover effect on a button, and here's a cool loading transition state. And those things are definitely cool. I don't consider that design engineering, because you could achieve most of those things by just telling an agent to do them without looking at a single piece of code. To me, that's just very advanced, sophisticated motion design and visual design. And that's cool, but I don't consider that design engineering, because I don't think that's a full job for you to do as a career.
But design engineering now, the people who I think of as good design engineers, do the kind of thing I was talking about before, where you step much more into the engineering side of work. So you're still a designer. You're still caring about product nouns and verbs and the visual design, but then you really work with engineers closely, and/or are directly involved in implementing and writing code yourself.
And you fully understand, well, not fully, you don't have to be full full-stack, but you have a deep understanding of the technical architecture of the product you're building. You really are like, okay, given the shape of the backend data, what is possible in the interface? That sort of design work. Given what models are capable of and what kind of custom skills we're building into this product, how do I explain to users what the capabilities of this product are? Truly digging into the technicals, and not living in the world of just user interviews and the market and what color is the sidebar, but caring much more and working much more closely with the engineers, and almost always, I feel, implementing a lot of it yourself.
I've always done my own front-end work just because it's easier. And then the engineers I work with are usually thrilled because they hate CSS. And then they get to go work on the more interesting, difficult stuff. I'd say it's the syncing, or the back of the front-end work, you know, the more logic stuff. They really get to engage in that, and they don't have to worry about whether this is the right border radius on something. I don't think they... they very rarely care, and I do care. So it's always worked out well to kind of do that. So I just define it as someone who's deeply into the engineering side of the design.
In your past positions, current positions, through friends who are also designers, what have you seen great engineering and design collaboration look like? Be that with designers or design engineers?
I mean, I do find, being a design engineer, I've always had amazing collaborations with engineers because, again, you do the bit they didn't ever want to do. And you take away what I assume is most of the tension between designers and engineers, which I've never experienced because I've always been a front-end person. But I've heard of people, you know, there being the tension where someone's made some perfect Figma file and the engineer hasn't implemented it exactly to spec.
Yes, I've been there. Yeah.
Yeah. And obviously there's a tension there, because someone made something
And immediately, there's a nice gradient, and it's mobile, and doing gradients would absolutely wreck the scroll performance. And again, we're talking about years back when this was a thing. Yeah. And now you're going back like, can you just do a single color? Or you try to explain what you can do with either your performance limitations, or think about low-end devices on Android, or iOS versus Android, where a designer might be more into the iOS world. And again, a good designer would not, but oftentimes the engineer would come and say, all right, we have constraints here
Right.
for whatever reason.
Yeah, I think this gets into... I don't want to blame the designers, because I think it's a failure of tools. The design tools that everyone used, Sketch and Figma, these pixel mockups, they have no relationship to the constraints of the medium, which is whatever you're building for: iOS or desktop or the web. You have to understand the materials you're building with in any design role. Yes, a table designer would not be oblivious to how oak performs in certain contexts, right? Or how pine dents in a certain way. And in the same way, I think if you're designing for the web and you don't understand performance, or how your app is fetching data, and what's the loading time, and what if there's race conditions, I think if you're oblivious to that, you'll end up with bad design solutions and then with really bad relationships.
Yeah. So I'm hoping agents actually help solve this, right? Because really, designers using agents can step more into the engineering side and also use agents as a coach and a learning tool and say, "Okay, the engineers came back and told me that we have this error state I've never heard of in my life. Explain to me why that error state would occur, and help me build a state machine," right? There's all these new tools available. But those old frictions, I have to assume, were just tool-based things. And I would say designers that are maybe stuck in the past, who don't want to let go of, oh, I make pixel mockups... I don't know if anyone's there anymore, but if they still are, I can't imagine that relationship going well in the future.
And then you mentioned Figma. Figma was such a popular tool for a while that it was the interface between engineers and designers. How do you use Figma today, or how have you used it in the past?
I think I've always used it not to get to high-fidelity mockups at all. I think to get the rough shape of things past a notebook. In a notebook you can be like, okay, the shape of it is like this, but it's such a rough sketch. And Figma is useful, I find, for being like, okay, exactly what color is the right amount of contrast to draw attention to this element on the page, or exactly what size does this text need to be to flow well. But I would never take it too high fidelity, because of course everything always looks different. I mean, I primarily design for the web. Everything looks different in the browser, right? It depends on the font rendering and all this kind of stuff, and responsiveness, and when exactly your breakpoints are. So I would always take it into browsers pretty early on. You get medium fidelity in Figma, the shape of it, and then you take it into a prototype and then you can really tweak and refine.
And of course with agents now, I just point an agent at my Figma mockup and I'm like, just get all of that in there, and then make what's called a jig, which is where you get little sliders and variables attached to a little... a jig.
A jig comes from woodworking, where you make a little device that helps you do one specific job. So with design, when you're working on a live prototype, you say, okay, here are the variables I'm not sure about. I'm not sure about my headline size. I'm not sure about these colors. Give me sliders and color pickers and all these things. Or the animation curve I'm not sure about. Give me, and then I'll, and then
I'll tweak it live and then when I have the values just right in the live version, then we'll commit those to be the actual values. Much faster. It's like build your own Figma as needed.
And inside of GitHub, how do designers work? How do yourself and other designers work? Cuz it's a bigger organization. You're building a tool for developers with a bunch of different parts.
I mean, I don't spend that much time with the GitHub design org proper. I know some people in there, but no.
Because you're Next.
So, the GitHub Next team is a little bit isolated. Like, we still have good relationships with the rest of the org, but a little bit by design, we're supposed to be kind of like the R&D team out on the side, kind of doing weird stuff and then trying to convince the rest of the org that we're right and they should pay attention and do our thing, which is a whole different politics thing. So, I kind of see the designers, but I think they have a very different job to me, in that they have to uphold quality standards and they're working on a big design system and an old Ruby on Rails app. They've got very different constraints to me.
Yeah. And they have millions or even more, like tens of millions of customers who... that's now a different thing. They've been used to certain things or finding certain things or certain patterns.
Exactly. Exactly. Changing anything on GitHub.com is a whole political thing, because customers expect, you know, to come in and they have critical workflows; you can't move their button. But the team I work on, there's me and one other guy who are both design engineers, so we both do engineering work, but we are more the design people, and everyone else is much more engineer-like, and they're all really senior, so it's fun to work with them, cuz they've seen everything, they know how to architect an app. But we don't work that differently to engineers, I find. I mean, we're putting in PRs the same as them, we're working in the same tools and materials. It's much more that thing of what bits of the app we care about. I of course care about the interface part, and then my engineer partners of course care about the back end and the data flow and everything. So same tools but different considerations, I guess.
I wanted to take some time to mention our presenting sponsor Turbopuffer. Turbopuffer couldn't be a better sponsor for this episode, as they have what I think is one of the most refreshing brands in tech right now. Turbopuffer's co-founder Simon describes their brand as hardcore and whimsical, and their design perfectly reflects this. Turbopuffer the database is extremely hardcore AI infrastructure. It supports critical search workloads for Anthropic, Notion, Bridgewater, Cognition, and many more of the largest and fastest growing AI companies in the world. They take reliability, performance, and scalability very seriously. But Turbopuffer, the team, refuses to be another sterile, boring enterprise tech company. Just look at their website. It's exactly what you would want the website for a database to look like. It's simple, modest, and transparent. They even put the limits on the homepage, but it also isn't boring. It's nice to look at. It's very well thought out, and their little ASCII diagrams perfectly blend hardcore engineering with whimsical design. Plus, if you've ever tried to generate a good ASCII diagram with AI, you'll quickly realize that they create all of these by hand, not with AI. Turbopuffer's brand is the perfect reflection of who they are, and that's why it works so well. The brand is hardcore and whimsical, and so is the team. I highly recommend Turbopuffer not only as the extremely scalable search engine for AI, but also as a group of people that will go to exceptional lengths to deliver for their customers. You can check them out at turbopuffer.com/pragmatic.
I also want to mention our season sponsor Entire. Maggie mentioned that as a design engineer, she's still putting in PRs the same as everyone else. And every one of those PRs is pushed to Git hosting, but there is a problem with Git and agents. Git is increasingly becoming a bottleneck for modern agent-heavy software development. Devs are creating more code with agents. These agents are pushing more code, and many devs are running parallel agents that are pushing even more code. GitHub is clearly struggling to keep up and has frequent outages. So what is the solution? Entire was founded by GitHub's last CEO Thomas Dohmke, and here we built Git hosting for agents from scratch. Entire was built to be very fast and have your repos regionally close to you to reduce latency, allowing for fleets of agents to push in parallel. And Entire's performance is next level. The repo can handle 418 pushes per second, and that's up to 89 times faster than every other competitor on the market. When GitHub is down, you can still keep working, and you don't need to migrate away from GitHub. You just sign up to Entire and the platform mirrors your repo. Oh, and one more thing. Have you ever wondered what prompt resulted in this code being generated? I find that the prompt and conversation with the agents carries more information than the PR itself. Entire captures all the prompt history with your agent right in the repo. Easy to check back. And it's got a pretty innovative UI to show all of this. If you're looking for Git hosting that works even when GitHub is down, head to entire.io/pragmatic, install the CLI, and mirror your repo with a click. I've already done it. Oh, and did I mention that it works with any agent and is open source?
And then with LLMs, since they came out, how has your design process changed? You've already mentioned how it's now a lot easier. For example, if you have an interface design, you can ask Claude or Codex or any other agent to implement it. But what has changed between the notebook, between your mockups, between the prototype that gets there? What parts are easier and maybe what parts are trickier?
It's definitely all faster and it definitely all feels a bit easier. I will say, I definitely remember suffering through making elaborate prototypes in Figma, because that was going to be faster than building more complex things in the web, and then trying to show them to users with these Figma clickthrough prototypes that obviously everything is faked.
Oh, I remember that, on the user thing, and then sometimes you have to have these weird things that click here and then...
So you're in these user interviews and the user's trying to click something that you haven't hooked up to a fake new screen, and the whole thing just feels a bit insincere. And then even when you did make prototypes, you couldn't spend any time making them look nice. So then sometimes the users would get a bit distracted or confused by the fact it looked terrible. But you're like, "Yeah, but we only had a few days and then we've got to move on to the next prototype." And now I wouldn't have that problem at all, right? Now I can build something that really looks quite sophisticated, and not with the agents doing it on their own. They need a lot of direction, I find, to make something that I would think was an acceptable interface.
But it's so much faster. Sometimes I have even mocked up quite high fidelity stuff in Figma and then just been like, "Hey agent, here's the Figma file. You know, goal mode. Use Playwright, take screenshots. If it doesn't look like the Figma, keep going. Keep looping until it looks exactly like the Figma," and it'll get there. And so that was just, I can run that overnight. I don't have to do anything. In the morning, the interface is pretty much there as I specced out.
And to you, I guess this was grunt work, right?
Oh, totally. It's totally, cuz you're like, this isn't complex. This is like, okay, again we're building another sidebar, right? Let's get the React components, and like, but...
And you don't really care about the code quality or tech debt or any of that, like it's going to be thrown away anyway, right? Like if it works, because of that, then you'll build it, and if it doesn't work, like, great.
Doesn't matter. So prototyping is a whole different game now. I feel like you can prototype much more ambitious things. And then I do love the whole jig thing. It feels like this malleable software thing where you can make the software whatever you want it to be, where I'm not constrained by Figma's available tools. I can be like, okay, I'm trying to make, like, this animated constellation map was something I recently did. What speed are the stars moving at? What zoom level are we at?
You did that?
Yeah. And like, what degree is the gradient at on the screen? And you just put these all on little sliders and you just tweak it, and then you think, "Oh, I really want a slight green grain over this." Tell the agent. You know, it honestly feels like magic a lot of days.
Do you have that?
Oh, I might be able to pull it up. Yeah. Let's see. So, this is something I built for the new GitHub Next website. We have a kind of basic one up, but I wanted something a bit fancier. We want to make it a little bit easier to publish research frequently. So, I was playing around with some jazzy homepages. And this whole thing, this is like an animated spinning constellation. And these are some projects that the team's worked on over the years.
Amazing.
And this kind of thing, to prototype this, right? Of course, I started with a sketch in a notebook, but then made a bunch of jigs where I was like, okay, what color's the background, and how fast is this map spinning? And then there's a bunch of cool physics on the stars that I didn't have to do any physics for, you know, I didn't have to know anything about that. You know, how big are these when you hover over them? And then I did this cool scroll effect down into the proper page. And all of that, right? I didn't write any code. I made a bunch of different jigs and tweaked and adjusted.
Exactly. And so these were just, imagine, things on the screen. For example, the color, you could go from blue, green, red, and then you figure out, all right, you like this the most, the sizes, all of those. Right.
Exactly. I can be like, make me a slider for controlling the animation variables on the constellation experiment. So, we'll let that work. But that's the kind of... it's really simple. If you tell it, make me controls or make me a slider for whatever it is you're trying to do, just think about, okay, what are the ways I need to tweak this to make it work? It's pretty good about just making them. I mean, I have a skill so it knows how to make them so that they look decent, right?
And based on your preferences, you know, what you figured out works.
Yeah. Yeah. But then it's great. It just hooks it up and then you can just kind of play with it, and it's a whole different way to design. It's very, you know, Bret Victor, the live programming stuff. Bret Victor, he's an engineer, also a designer. I mean, he's off doing some crazy stuff now. He did a set of talks between, I don't know, it was like 2010, 2013. One of them's "Stop Drawing Dead Fish." And it's about how programming is not a very live medium. Like, you know, you code in the editor, you do your whole build, and then you look at it in the browser or wherever it is. And these two things feel very disconnected, and it's very hard to...
You're tweaking variables in code, and then you have to go see the effect. And it's not a direct connection.
And so his whole thing was you need to have instant direct feedback at all times to be able to look at the artifact you're making and directly tweak it. And it's like this kind of stuff come to life. Now we finally can do it, but before you couldn't do that. There was no programming system in existence that could give you a live preview of every possible variable or thing. The closest we have is dev tools in the browser.
Well, another topic which is very related to this one, while we're waiting for it. One thing with AI, of course, you're saying it's easier to build stuff, but with AI it's also easier to design. When I was asking people how they work with designers, an engineer replied something which I hear a lot more, which is, I'll quote this person, Amir: "As an engineer, I ask Claude Design to generate 20 high fidelity alternatives, and I keep iterating till I land at a perfect design. It's very enjoyable and feels much faster than interacting with a design team." So it is a thing, saying, oh, I have these tools, I don't necessarily need a designer. I can just ask for all these, and in this case it will just give me 20 alternatives. I'll choose the perfect one and boom, I'm done. As someone who is a designer, what's your take on this? And by the way, this is not the first time. I think there's always the thing, do we need designers? Do we, as engineers, need product managers? And of course, product managers will also ask, do we need engineers, and so on? But what do you think? Where could this be valid? What could people be missing?
I think in cases where you don't have a designer to hand and you're trying to validate a product or prove a hypothesis, but you need an interface for it, it's great to just use a model to be like, sure, make me 20 high fidelity designs. I'm sure I might look at those designs and be like, well, it all depends on taste too, right? I might look at them and think, okay, well, these look obviously generated by AI to me, and I don't know that they're going to do the job that they need to do. Again, it depends on the context. If it's something simple, needs a button and a sidebar, fine. But if it's, you know, we have some new primitive we're trying to figure out the shape of, that's when you really need a designer to come in, and really it's just someone who's assigned to think through the problem properly, put in the brain work, and be like, okay, do the experiments, show them to users, figure out what people actually understand and don't. I think that's really the labor a designer should come in and do. But I think it's completely fine for developers without access to one to use models as much as they can.
But when models do designs for me, I of course just look at them and think that is terrible quality. That is awful. Even to the extent that I think they've been prompted with what we would call universal design principles, but they sort of don't understand nuance and context. And then I think they stick to them too strictly. So I find the agents want to put a label on everything in the interface.
Label meaning?
Like a small bit of text. There might be a little button to close the sidebar, and we would usually use an icon button for that, because most people are trained that this little button near the sidebar means it'll close it, and they'll experiment with it and they'll click it and they'll figure out that is correctly what it does. But the agent will write "close sidebar" or "close modal" in the top right hand of the modal, and you're like, there's now a lot of text on this page. And they'll just put four lines of instructional text over a button, and just things where you're like, I understand in the model's mind it's thinking, oh, this is how we make the interface explainable to the human, you know, but actually it's really bad design.
Now, talking about the UX and UI for AI, what do you think good UX and UI looks like? And I know this is a bigger question, but so far what have you learned of what works and
what doesn't? I think when it comes to this, like, what is the ideal UI of AI, which is a very big question I don't know that we'll figure out the answer to for a while, but it's something like some things haven't changed, like in terms of what is good interface design for software. There's still a lot of principles here, like what formats do we have to show users data, you know, it's like dashboards and sidebars and documents. Like there's a lot of things that haven't changed.
We actually had a good story about this from Elicit, where the app was helping scientists extract data from papers, and they were very used to doing this in Excel spreadsheets. They'd get a big Excel spreadsheet, put the papers one on each row, and then they'd extract each piece of data into the cells, right? Pretty standard process. But of course in the beginning of Elicit we were like, oh, that's so old school. There must be some much better interface for, you know, seeing the data extracted from papers. So we tried all this crazy stuff. There was like infinite canvases with cards spread everywhere. There was one that's like a bunch of cards that are all linearly stacked. We were like, maybe it's more like Notion with these composable documents with rich interfaces. Like we tried all this stuff, and every user interview I did, users were just like, this is very confusing. Can I just have a table? Like every time.
And so after a couple months of this, like, oh, what's the new UI of AI? We were like, oh, it's a table. Well, [laughter] at least in our use case it turned out the interface people were using was the best because it was the most familiar to them and it caused the least amount of cognitive load for them to use the tool. So we were like, right, back to tables, that's fine. Like it was good to go on that journey. But sometimes we have this notion of like the UI of AI will be so wildly different from our current imaginings. And sometimes it's really not at all. It's like you can do more powerful things. I think there's lots of leverage, but I expect it to all be built with documents and sidebars and cards and tables and all the classic things we've worked out do work and people know how to use.
Yeah. It's interesting. Yeah. There's a cognitive load in learning any new UX UI paradigm. It can be confusing if your existing tools don't work like that, because again people will still use operating systems and Google Sheets or Excel or Office documents, etc. And even if a small subset of people are, I don't know, familiar with this new UI stuff. You remember probably parents or grandparents, teaching them how to use the phones, the interfaces, and once they learned it they're good. But then do you really want to do that again?
Yeah. No. [clears throat]
It's interesting. [laughter] Yeah. So there will be something similarish.
Yeah. It's kind of like how we started with chatbots, right? It's like chatbots were not the final form. Like if we could look at Codex, like there's a lot going on here. Sure, there's a chat window here, but there's all kinds of things going on with git worktrees over here, and then we've got this whole interface, and I can do annotations and debug in this. Like there's a lot of extra stuff where you take the familiar primitive and you expand outwards from it. And we did similar things at Elicit, like there was still a table, but then we baked a lot of stuff into the table and around it. But you start with familiar primitives. So here I can like
we have our jig ready.
Well, yeah, let's expand this. So here's a jig. This is using dial kit. This is built with this. So I can change the speed of the animation to see like, oh, is that too fast? Like is that too slow? The intensity is like how much gravity is being applied to the stars here.
Oh wow.
How much variation is there in the stars? See, this is sizing and scroll. So this is like when I scroll down to have this zoom out, like what's the resting height? I think this is, when this is here, like the resting height of this, resting logo size. Let's see, what's this. This is like, I have some glass effects on the surface, it's like, oh, it's on the top thing like that. How much is that glass effect? But see, that's too bright. Then you can't read the words. All the way down here, you can barely tell it's glass. So there's some sort of happy middle ground you want to figure out, like how much contrast between the background. So this is where
You weren't kidding that this is like your own almost personal Figma, if you will, or like design tool.
Yeah. And then because it's really difficult to mock up static versions of this in Figma, which I have done. Yeah. Like for a minute we were thinking maybe like a gray background, trying to figure out how much grain and stuff. Maybe like a wavy one. Yeah. And then I'm trying to figure out here, like, okay, is it halfway off the screen? Maybe there's like a slice out of it here. So you can play around with layouts and ideas, but you can't get a sense until it's live in the browser.
Yeah.
Of how this is going to feel and how much scroll you need to do here. And then like the themes and the colors and textures and everything. So I work with this a lot now. Whenever I'm designing something visual, it's so much easier to have live variables, which again would not have been possible before. You could have built this by hand, I guess, in the olden days. It just would have been practical.
Just coding the thing itself was slow enough.
One interesting thing you talked about is capability gaslighting. What is capability gaslighting?
Okay. So I think this is similar to, like, Ethan Mollick has this phrase, the jagged frontier. It's pretty much that, but I came up with capabilities gaslighting before I read his work. But so that people are familiar with the concept, it's that models are really, really good at some things and really bad at others. And it's really hard to predict, when you give it a certain task, which of those it's going to fall into. And then capabilities gaslighting was the feeling I got early on from using them, where they convince you they're so capable because they'll really impress you on one task, and then you try them on something else and they fail, and you kind of feel like, I feel like you kind of gaslit me, like imagining that you're this extremely intelligent agent or model, and then you fall on your face. And then sometimes I feel like they fail but I haven't totally noticed how badly they've failed, because I still have this belief that like, oh, but you're a frontier model, like Opus could never really get this wrong, and it totally does all the time.
So it's such an inconsistent experience. And it's so different to working with a human. Like if you find a human who you feel is really an expert in a topic and you work with them, it's very unusual for them to be inconsistent in their performance, right? Like it's very rare for them to suddenly forget all of their expert knowledge on a topic. And if they did, you would be very like, are you having a mental breakdown? Like this is very weird behavior. But this is how models behave every day. Like one day they will actually perform really well on a task, and they could fail at that same task the next day because you prompted differently, or they had different context available, or, like, random, whatever it is, stochastic outputs, they just didn't do as well. So I think it's hard to work with them because you never quite know what you're going to get.
Yeah. I guess we just need to keep this in mind, right? Because this is ongoing. Like of course we always say, and we have seen, that the models improve, but you want to be skeptical.
Yeah. Yeah. Of course there's some baseline with frontier models, like they're not going to perform below a certain bit. But, you know, of course the models are changing all the time, and people complain about this, that you upgrade to a new model and it changes what it's good at, or it doesn't perform to your expectations on certain things. It's like it's hard to predict.
You did a talk titled "One Developer, Two Dozen Agents, Zero Alignment: Why We Need Collaborative Engineering." Can you talk about just this observation, like one developer, two dozen agents, zero alignment?
Yeah. So this is a problem that the GitHub Next team has been focused on, trying to find various ways to solve, I think for well over a year now. It's this acknowledgement that we are now working with these agents locally on machines, and it speeds up each individual person. Like us alone with the agent, we can go really fast. But software is always built on a team, right? You're always trying to align with the product managers and the designers and the other engineers on what decisions you're all making, and we don't actually have good tools in place to do that. It's sort of like we have Slack, people might have something like Linear or GitHub Issues where they're doing issues, but there's tons of pre-planning work before you write an issue, because usually when you write an issue, you're ready to hand it off to an agent. But before that point, you have to have agreed upon your approach, and like, should we build this feature? Is this feature the right shape? Does it have the right interface? Do we have some sort of database migration we have to do? There's all this upfront work and not many good tools to do that work in, is how we felt, especially not ones with agents involved in them.
So the biggest gap seems to be like we have agentic coding tools, but none of them are real-time multiplayer. All your sessions are private to you, often locally on your machine, so you couldn't possibly share it with someone. This is starting to change now. We have seen some products come out that are trying to do live multiplayer agent work, like Buzz from Jack Dorsey, and ACE, the prototype that we made at GitHub Next, was in this direction. It was like shared compute, shared sandboxes in a Slack-like interface. So you're coding and talking at the same time.
Mhm.
So people, I think, are beginning to realize this is the next big thing we need to push on to improve in our agentic coding tools, but it's still totally unsolved. Like even once you have a Slack interface and there's an agent in there, you still need tools to agree on things. This is where I'm kind of pushing on, like maybe decisions need to become a sort of first-class primitive. Like when an agent presents a decision to a human, you first need to have the decision be large enough on screen to give you the information about it. But then also you probably need someone else in your team to come help you make that decision, or at least give input on it. And then someone has to be responsible for having made that decision. You need a record of like, here's the things we decided, here's the context we're going to feed to the agent. You know, there's some moment where we've written a clear set of issues or specs that we're going to hand off to be implemented, and then some verification check on the other side.
But you all need to be so aligned up to the point of implementation, because in the old world implementation took so long you could adjust along the way. But now, because there's a sort of hard handover point to an agent, you need to all be aligned up front in a way that we're not at the moment. Like even our team, we feel this all the time. It's like we're all individually running on our machines and it's hard to stay aligned on what we're doing.
And you mentioned GitHub Next and the prototyping that you're doing. Can you talk a little bit about the types of projects you've built, that you've experimented with, some learnings that you've had? Maybe give us a direction of things you're now excited about exploring.
Yeah. Yeah. So I joined the team a little less than a year ago, but most of the time we were working on this prototype ACE, which is like Slack plus cloud compute, like sandboxes, so microVMs, and a bunch of other stuff in there. Like you can [clears throat] open PRs and review your code and that kind of stuff. So it was kind of an all-in-one multiplayer workspace. And it was a really useful prototype, but it turned out to be extremely ambitious for our small team. We had like three or four people working on it at a time. And we wanted at some point to take it to a real product, but it just turned out that was not feasible with the number of people we had. Yeah.
But we're trying to take bits of it and ship it to the rest of GitHub. Like sandboxes have now gone out to be part of the GitHub desktop app. That's being used in other ways. And I think it's also made leadership take the multiplayer thing more seriously and try and find other ways we might ship that to the product.
And this is a really promising direction. I mean, we already have a deep dive out, by the time this podcast is out, about Ramp, who have built just this inside, and they made it collaborative. It's forced collaborative. You cannot make anything private. Oh yeah, they have multiple interfaces. They have this running, it's called Inspect, and it runs in Slack. There's a Chrome plugin as well. There's a web interface. And they did find that in the Slack channels, with product managers and designers, they can talk to the agent, Inspect gets all this input, and then for all the sessions, anyone can join a session and anyone can prompt and change the direction. Which actually was a really concerning point; they weren't sure if they wanted this because of people's privacy, but they were like, well, it's just civilized, so it turned out to not be an issue. But because of this it spreads better and it has all these feedback loops. And the reason it works really well for them is it's integrated this agent to all of their internal systems.
And their cloud machines are like a full-blown developer machine, which is a lot of work, a bit like a cloud development environment. Which is all to say, doing this for one company, specifically for them, it's possible. Doing it as a generic solution, probably more difficult, but I'm sure it will come, or I'm sure there will be attempts. But, you know, this is the difference between scratching your own itch and doing something that works for so many. I mean, Slack, right? They launched their whole developer experience IDE thing, which, like, they might be able to make this experience happen, or like Buzz from Jack Dorsey is kind of going in this direction too. Like there's clearly lots of people realizing this is a problem.
And you're also experimenting in this direction, clearly.
Yes. Yeah. Yeah. Yeah. So one of the things, like when we tried to build this whole thing, it was too ambitious to do it with all the microVMs. But now we're kind of scaling back and thinking, okay, what bits of this can we prototype in ways that would help GitHub leadership have a bit more direction of where they should run next? Like, you know, one of the functions of our team is just to run out ahead of product. Product knows the things they should probably do in the next six months to one year; we are much more like, what is a big crazy swing GitHub can make, or what's more risky, or what's the far future. And so that's why we're trying to prototype. Okay, people are now on this multiplayer sessions thing, what comes after that? So now we're interested in how do you make proactive agents that aren't annoying, is one of the things we're trying to explore.
Because proactive agents, you think, we have this dream of like, okay, assume intelligence is free or so cheap that it doesn't matter, right? That's a fun assumption to make. If you could have 100
background agents running, what would be the least overwhelming, least annoying way to consume any of their outputs is like an open research question we're trying to figure out. How would you have agents working in a document with humans where it's not, again, overwhelming or annoying for them to be making changes or doing helpful background work in a way that doesn't feel invasive or interruptive? So we try and take these kind of design questions and figure out prototypes that could solve them in the hopes that, you know, if one day GitHub wanted to build something in this direction, we can then be like, hey, here's all the research we did.
A question I wanted to ask you, and it's coming from a developer angle, is what happens, in your experience, observations, when we start to give more and more decisions to the agent. And I'm going to quote a developer, Jorge Manrubia, who wrote an article called "Oh my craft." He wrote: "The need to intervene on the small stuff is decaying quickly as models improve, and when models are capable enough to handle those details themselves, spending too much time on these details starts to feel like a poor use of human attention." And he writes about how he really doesn't open his code editor anymore. It's been months since he wrote a single line himself. But the thing is, he used to live inside and see all the details. And there is this sense that craft has to do with being there with the details. And of course you're coming from a design perspective, but what have you observed of your own craft changing, decaying, or, you know, what you're seeing? Do you think there's a danger, as there's this temptation, as we talked about, to just, good, make this decision? We just did this with the slider, which didn't matter.
Yeah.
But it can matter.
But it can matter.
But does it matter?
I mean, it's hard because I'm not, I would say, a true engineer in the sense that I don't necessarily care about clean code that much, right?
But you care about great design.
But I do care about design. But the thing is, they definitely can't do design to my standards yet. I have to still be really involved to get the design to look and feel the way I would make it. So I'm annoyed almost that I'm always in there being like, "No, that is a terrible transition. Here's how we should do it," you know, or like this.
So you're in the details. You're not letting go of those details.
Yeah. Because I wish I could, I mean, I've written—
Do you design— well, I don't know. I want to push you on this.
Yeah. I keep thinking this: I keep trying to write design skills that tell the agents exactly my design preferences, which won't be universal. Like, I have preferences about how much padding I like between a border and an element, right? And that's not everyone's preference, and it's not right for every product. But for me, there are set rules. It's this much padding, it's this kind of border shadow. It's pretty standard. And I've tried to write these rules, but then they don't universally apply them properly and it doesn't work. And so I always end up in there changing specific values that you would think by this point in time would have been automated away, the way that we all talk about agents taking over stuff. You would think it would be so trivial for them to implement a certain kind of, whatever it is, opacity level on a border.
So I kind of think, if I woke up tomorrow and the skill just worked, or the agents just had enough context and they just designed exactly to my specs, I wonder if I would be like, oh, all the fun bit is gone. Is it satisfying to have a gorgeous interface appear in front of me, but I didn't do anything to make it? I'm not sure. I also think about this with interior design stuff. Have you been on Pinterest anytime in the last two years?
I'm familiar with it, and my wife is. When we were decorating a new house, I remember collecting these things. We worked with an interior designer, and my wife collected a lot of things, but I also was there, and you have these kind of, we're kind of thinking like this but not quite, but when you mix this, trying to explain. We didn't have the tools beyond, you just take screenshots, so, yes, don't know what print, I guess.
You know, a lot of it is AI now if you get on there.
Oh, really?
There's gorgeous rooms if you search for certain things, but you look and you can kind of start to see, like, oh, that's for sure AI, that's for sure AI. Which is definitely problematic when it's of a physical space, because then you go, well, the light might not be accurate, or this is all just fake. And they look beautiful, but you kind of think, well, someone didn't make that room, and that's not a real room that exists, so it's almost irrelevant to my interior design.
And I kind of wonder if it's like this with interfaces, where it all objectively looks gorgeous because it's been trained on very much everything that was on the web that we've clicked thumbs up on. But with interfaces, if they all become extremely slick, like, take what's popular right now, the Linear and the Vercel aesthetic, right? Very minimalist, very clean, very white, a little bit of rounded corners. If all the agents can implement that, no problem, just dead on, people will start to design in different ways, because that will become a tell that you've just used an agent to do the design. And if you want to stand out, you need to have some human do something new and different.
I think it gets into design, although it has universal basics, like is the text big enough to read? Do you have enough space around things? A lot of it is fashion. It's in fashion right now to look a little bit like Linear. It's not in fashion to look a bit like MySpace, but it was a while ago. And 10 years in the future, if you look like Linear, you're very out of date. You look like some old crafty piece of software, and there'll be some new aesthetic that comes up. So the models can't necessarily understand that design is within a cultural context, and that cultural context is always changing, and it signals different things to people. And if people read it as, oh, they didn't care about this site because it's just the bland basic code. Like, Claude has a specific design language. It's very noticeable, right? Cream background, slightly red text, there's eyebrow text on things. You look at it and you're like, "Yeah, Claude generated that. No human was involved in this. I don't know if I'm going to bother looking at it." So it gets into this: the aesthetic style you pick communicates to the person viewing the product or page. And things that clearly were made by agents, I think we might start to reject out of hand.
Maybe I— I have this thing where when it's written text, I can immediately tell if it's AI-written.
Oh, I cannot.
But I can never tell if it's because I write a lot. I also read a lot. I read a lot of fiction. And so it just feels off immediately. And it's not just Claude. It's the repetitiveness, the adjectives. But I can never tell for sure, like, am I the one who notices? Because you said, "Oh, you see it's Claude." I'm not sure if I looked at a web page that I would see that it's Claude, because you're so into this the same way I'm so into words or nonfiction writing.
Yeah.
So I do wonder if maybe, as an expert, you can always tell, but the experts are always a small percentage of the population. We don't know. But I do agree with you that it is predictable in ways that humans are luckily not, I guess.
Yeah. I definitely noticed the writing one too, because I'm the same, I'm a writer. I love writing, and the minute I hit any sentence that is remotely like, yeah, just close tab, never mind. And even going back and reading nonfiction books that were written not that long ago, you go back and read a book written 50 years ago and just the opening is so refreshing. You're like, oh, these are such novel words. There's none of this "if this, then that" or "it's not X, it's Y." There's not a single bit of it to be seen. You just breathe in original human writing, and it's so refreshing that you go, yeah, I can't stand to read anything that's agent-contrived.
Speaking of original writing and human writing, you talked about the concept of a digital garden. What is a digital garden?
It's a blog but with some extra rules attached. So it's a blog where every piece you put up does not have to be finished, as long as you clearly signal to the audience it's not. So it's a blog that you grow over time. You can put a piece up that's half done and then update it later; you have an updated date. And I communicate it with three different stages my posts go through. So I have seedlings, budding, and evergreen, very much in the gardening metaphor. So I can put up something that is a half-finished thought and mark it as a seedling, and then come back to it later and finish it up or fix it.
Writing in this way has allowed me to publish much more than I ever would have otherwise. I have perfectionistic tendencies. This was very much a counter to that, where of course I want some of my stuff to be my best work, my most polished, but it's completely unrealistic. I would never put anything up if that were the case. So working in this way really freed me. I mean, I started my garden in 2020, I think. So it's not been that long, but I've written a fair amount, more than most people have blog posts on their website, because some of them are frankly three paragraphs long.
Yeah. Or some of them, I have big long essays that have taken me a long time to write, but halfway through the essay you'll hit something that says "draft in progress," and everything below that point is pretty rough. One of these I started three years ago; every time I have some free time I go back to it, I finish one more paragraph, I move the little draft notification down. It allows me to work in public, which I also think helps people get to see behind the garage door a little bit, because often the notes below the draft point are bullet points or like, oh, I should say something about X in here. I'll just kind of put it on the website. As long as it's marked appropriately, I think it's fine. It's about this contract between you and the reader. As long as you're communicating to them that these are rough notes I haven't finished, that's okay. As long as you're not saying this is my best finished work, I feel like people read it in the right spirit.
Somewhat related to this, again just going away from AI, you also talked about barefoot developers and home-cooked software.
Yeah.
What are barefoot developers?
So this was, I'm not sure I remember what year this is. It might be 2024, which feels like an eternity ago. It was not that long ago. This was at the Local-First Conference in Berlin in 2024. I did this talk, yeah, home-cooked software and barefoot developers. Home-cooked software is a phrase that comes from Robin Sloan, which is software that you make for you and your family in the way you would make a home-cooked meal. Yeah. So it's about this app that he built for his family where they send each other little video notes, and it's not on anyone else's server. He doesn't have to pay for it. It's not managed by some big company. And this is the way more software should be. There's lots of little pieces of software we can build for ourselves that benefit our lives and our families, and it doesn't have to be $2.99 on the App Store or sell my data in exchange for this software.
And I made this case at this conference that I expect to see an explosion of this with language models. Okay, the writing was already on the wall. I don't think it was that crazy of a prediction, but we were still more in the early days. We didn't have agents yet. But now, of course, it's totally happened. There's personal software everywhere. Everyone's building their own recipe manager and their own household app and their own gym app, and there's all these great use cases where people can just make the software be what they want without having to download a bog-standard version of it.
But I extended this concept by saying, okay, home-cooked software is for you and your family. But there's another concept that I call barefoot developers that comes from Mao's China. They had this thing called barefoot doctors, where they took peasants from rural villages and trained them up with basic medical training, like giving vaccinations, you know, administering antibiotics, and then distributed them out to the villages so that you would improve healthcare for everyone. And it was a wildly successful program, across the board improved health, as you would expect for people who can't get to a hospital. And I think we need the same concept for developers, because so far developers are extremely expensive as a profession and extremely skilled. But there's lots of people with lots of needs who would maybe need little software built for their allotment garden or their street. They have some problem they need solved, and Google Sheets doesn't really cut it. They can't really just do a Google Doc. They want something like managing a schedule or managing inventory. And either they would have to all raise money together and fund it and pay for some piece of software that doesn't quite fit their needs but does an okay job, or maybe trades their data for this. Or if they had someone who was, let's say, this barefoot developer type, who's kind of a power user, someone who might use Notion and Airtable in the old world, they can now just use an agent or language model and build them the software they need for not many tokens, pretty cheap. They don't have to read the code. As long as it works, as long as they verify, it's fine.
I mean, it reminds me a little bit of the webmasters.
Yeah.
So back, it was a long time ago, but in the mid '90s, late '90s, the webmasters were just power-user administrators who would administer, for example, a school's Windows installations. And they were not the highest paid. I'm not sure how they got trained, but at some point every school had a webmaster. Sometimes they were volunteers, sometimes they were paid. Companies had their webmasters, and they would often operate the web pages, which was way more involved than later, of course; that's why there was a role. And then patch things, and it was all kind of, you know—
They might forget to patch stuff or they might not know.
And there were forums, webmaster forums, where they helped each other. Because what you talked about, the China example, that was of course a government program with a goal, but this was just a grassroots movement.
Yeah. And I think that needs to happen again. We definitely have these people to some degree. There's a techie person in each community who people go to for help. But I think these people need more support and tools. And I mean that in the sense of building libraries or frameworks for them, because I think there's a whole bunch of people now coming in as vibe coders or whatever, right? Building apps without looking at the code. But you know they're hitting terrible security foundations, and they're probably handling the data wrong. They're going to lose their database at
some point. Like, they don't have good foundations to build on because they're just telling the agent, like, "Hey, make me an iOS app that does X," but they don't really understand engineering. You know, I would advocate for something that is — I think it should be local-first, because as a philosophy it matches up with the need. It's like data should be local. You can just sync it between machines. You don't really need a database in the cloud. There's no point to that. A little bit of data sovereignty. They own the data. No one else can buy it or trade it or get access to it.
I just feel there should be much stronger local-first frameworks that give you these good, solid principles in place, good security, good data persistence, and make it easy for these barefoot developers to build, maybe with some good interface primitives on top. It's maybe something like what the next version of Airtable would be. This is what I would hope.
What advice would you have for engineers to learn from designers when they have access to working with a designer? Learn from, and also collaborate better with?
Yeah, I think it's in the same way that I'm advocating that designers shouldn't be afraid to get involved in engineering now, and vice versa, right? In the same way that a designer can sit down with Codex or Claude or Copilot and be like, "Hey, explain the back end to me," just not the details, the shape of it, I think in the same way a developer can sit down with an agent and be like, "Okay, I need to design a sidebar. What are the principles of good sidebars?" "Okay, I have my blog. Teach me about typography." You now have a very patient tutor who's just going to sit there and be like, "Hey, this is where you would vary line height. Here's exactly how many characters you should fit on a line." They can just teach you, as long as you're willing to learn and you want to expand into the design skills.
I think before, it was very hard to learn these things. You had to be on a design team, especially product design stuff, where it's like, how do you interview users? How do you do a usability test? Not generally available information, really, except for some books. But now you sit down with an agent and you're like, "Hey, here's the data I've got. Give me some ideas for possible interfaces that would represent it." Well, they're decent at that conceptual stuff, even if they don't always get the interface right.
Going back to your anthropology roots, what are things that us engineers, people building software, could take inspiration from or learn from anthropology? Like methods, approaches, things that have really served you well?
Yeah. Because I feel like I look at everything a little bit through the lens of anthropology, or I try to. It's like a pair of glasses you can kind of put on, in the same way you can be like, "I'm going to look at this like an engineer." But if I look at things like an anthropologist, you think about what are the unspoken cultural rules going on in this interaction or this context or this problem. And especially in the context of software, sure, maybe you're building a database, but you're building it for users in a cultural context, right? They have assumptions about how they decide something is trustworthy. They have assumptions about how they decide that something is worth their time, or how a flow should go. And it's all culturally contained.
It depends also on how international your audience is and who you're designing for, but there are cultures where time flows from right to left, right? Or from up to down. But we assume in the West, you know, time flows left to right, but that's not universal everywhere. I think just reading a little bit of cultural anthropology and understanding how varied people's worldviews can be — like, some cultures don't see different colors, but categorize colors differently in a way that makes them see them differently. They might see blue and green as actually one unified color. And if you're designing an interface, not that that necessarily directly applies, but I think it's understanding there's a broad range of ways humans can interpret something. And, you know, if you're building tools for humans, there might be some way you could teach them to see the world differently through the thing you're building as well, and change their perspective on how they interpret reality.
And then in anthropology, you said that one thing the traditional anthropologists would do is live with natives and be embedded with them. Is that something that maybe now, as AI is, you know, making code a bit easier... As engineers, I think it's pretty clear that an engineer becomes more valuable the more they take on the other parts of the business, the more empathy, the understanding. Could it be just an idea, when you have the opportunity, to just, you know, embed yourself with customers?
Yeah. Become a customer. I mean, this used to be the thing, right? Amazon had their customer obsession, which also starts with understanding the customer, but I guess this might be very natural as someone who's an anthropologist, which is like:
Go talk to people, go and try to be one of them, whoever you're building with.
There's a good point that engineers are about to have more free time in a certain way. I know there's an infinite number of engineering problems to solve, and maybe we solve them better by having that extra time. But a big part of this is, yeah, you could expand into more of the design side, and past the interface is really the user research side, and that is: go fully understand the domain you're designing or building for, and what real-world context people are using it in. Like, when are they pulling out your app? In a factory? Are they in a tube?
I don't know, a little bit of understanding context of use is a big thing that user researchers do, and then understanding the moment when someone reaches for your product versus reaching for a different product. When do they decide you're the right solution? That definitely gets into the user research side, but that's a great thing for engineers to expand into, if that appeals.
And as closing, what books would you recommend, ones that you enjoyed reading, and why?
One of my favorites that I give to most people is called Addiction by Design. It's about people addicted to gambling machines in Las Vegas, but it's an anthropologist doing it. And she talks about both living among these people and their experiences, but also the machine designers. Like, how do you design a machine that is so addictive that someone sits at it for 12 hours straight? It's a really fascinating thing. And how are gambling casinos designed to have no windows, so there's no time around you as you sit at this machine?
It's by Natasha Daw. And I read it in university and I love it. It's a total mix of cultural anthropology, participant observation, and also machine design and engineering, and how do you design addictive systems. Which, of course, you read it and then you think about phones and Instagram, and you reflect a little bit on what are these systems we're building for people.
Maggie, thanks so much. This was very interesting.
Yeah, thanks for having me. Really fun.
This was a very visual episode, so I hope you were able to watch it over video, especially the parts where Maggie showed her sketches in a notebook. And this was one of the most interesting parts of this episode for me. Even as Maggie spends her days prototyping the future of agentic tools, she starts her design process with pen and paper. Her reasoning is that agents cannot see, but us humans can. And I have to wonder, is it only true for visual stuff?
I mean, when I think back to my most productive brainstorming sessions with engineers, it was usually in front of a whiteboard drawing stuff, where we drew out our ideas about the system and the components.
Another extremely cool trick from Maggie is how she uses AI to generate prototypes with sliders to change parameters and change the design. She calls it a build-your-own Figma. And we did a live demo with the constellation animation that she played around with. This was so darn cool. You get Figma-like tools dynamically built for you. It's also pretty incredible that we can do this with agents in a matter of minutes. Just wow.
Finally, I appreciated how Maggie was honest about how she feels about the design craft. It's cool that agents make her work faster, but she said she's not sure she'd love it if an agent just spit out a perfect design. And yet, agents are getting better at design. It feels to me that this is kind of how I feel about code. It's great that the agent can generate good code, but it was something I liked being good at, and I was good at it. Now I don't write the code, but it does feel more transactional. I feel a bit less connected to the code and the creation of the code itself.
For design, agents are not there yet, but I didn't sense that Maggie wants to let go of that designer in her. I wonder if one thing to learn from her for us devs is that we should also have a notebook to make sketches of ideas, not just UIs, but architectures and systems. I don't know, but I might give it a go. At least it will make me feel less that I'm dependent on agents, and I'll have some AI-free space, if you will. Thanks for listening, and let me know how you like this less conventional episode. Thanks, and I'll see you in the next
Article published · Updated
