Creatine, AI Doctors, and Sleep Trackers: Notes from the First Don't Die Podcast

Open on YouTube ↗
Overview

This first episode of the Don't Die podcast is a loose weekly roundtable. Bryan Johnson is joined by Kate Tolo, his co-founder at Don't Die, and Dr. Mike Mallin, his lead physician. Johnson says he talks with Tolo roughly 150 times a day and with Mallin five or six times a day, and that the goal of the show is to share the ongoing work of refining his protocol. The conversation covers four things: what emergency medicine looks like from the inside and how patients can advocate for themselves (including with ChatGPT), a new study questioning creatine's effect on muscle, how body awareness develops once you start measuring yourself, and Andrej Karpathy's personal comparison of sleep trackers.

18 min read

From the ER to "future-focused" medicine

The episode opens with some banter about plants. Tolo's leafy background turns out to be fake vines. Mallin admits his wife looks after their houseplants better than he does, because he waits until they wilt before watering them. Tolo jokes that this is the ER doctor in him: he triages the plants and only responds when death is imminent.

The joke leads into a more serious comparison. Johnson asks Mallin how a day in the emergency department compares to a day working on his protocol, and Mallin says the two "couldn't be any more different." He describes emergency medicine as "banging your head against the wall all day long." In his account, patients arrive with problems that stem from a culture that doesn't value health, and they show up at the end of a disease process asking to be fixed. His current work, by contrast, is about optimizing someone's health 20 years out rather than trying to reverse terrible trends. He says the ER "was tough on me."

Johnson asks how much of what arrives in the ER could have been prevented. Mallin's first answer is "at least 90-plus percent," which he then revises to about 80% for emergency medicine specifically. Acute events such as car crashes, appendicitis, gallbladder problems, and infections are hard to prevent. Chronic disease, though, is in his view the dominant issue across medicine, and much of it is preventable.

Routine versus judgment in emergency care

Johnson recounts injuring his hands years ago when a five-gallon glass water jug shattered while he was holding it. The glass lacerated both hands and severed a tendon in his left index finger. A nurse told him they see about one such injury a day, usually from scooter accidents. She mentioned a 17-year-old girl the day before who had badly injured her face in a scooter crash. Johnson says he now wants to tell every scooter rider to stop. While he was in the ER, he looked up the standard procedure for his injury and found it laid out as something like a 19-step process. He asked Mallin how much of ER work follows routine and how much is the physician's own call.

Mallin says a lot of medicine is routine, with established standards of practice. The "art" is in diagnosis. He compares treatment pathways to pre-planned Disney rides: the doctor's job is to put the patient in the right seat. A hand laceration like Johnson's is easy, and Mallin says he could walk him through it with his eyes closed. The hard cases are vague complaints like headache, belly pain, or feeling weak and dizzy, which could be produced by thousands of different conditions.

America versus New Zealand

Tolo says the ER in America feels like a war zone, with people who seem to be dying and not getting seen. She adds that she doesn't know whether that impression is accurate, and that it has been ten years since she lived in Australia, whose system she considers better. Mallin agrees that American ERs are chaos, to varying degrees depending on the hospital. He hasn't worked in Australia but has worked in emergency departments in New Zealand. The biggest difference he saw was the ratio of physicians and nurses to patients. On a regular US shift a doctor might see about 30 patients, compared to maybe 10 in New Zealand, roughly three times the workload in the same time. In the US, he says, you are "not as much putting out fires as you are just trying to stop the speed of the burn." Critical cases like severe trauma and cardiac arrest get full attention. Chronically ill patients are checked to make sure they aren't dying and are then sent back to their regular physicians.

"Throw your complaints into an LLM"

Tolo's mother recently went to the ER with an eye problem, and she asks Mallin how people should advocate for themselves inside a system with its own procedures. Mallin prefaces his answer by saying his ER friends will "kill" him. His advice is to enter your complaints into a large language model, ask what should be happening, and then make sure it happens. He acknowledges that doctors often dislike patients who Google their symptoms. He argues, however, that this kind of active self-advocacy matters because the system is broken. Staff are overwhelmed and likely burned out, so a passive patient will be placed on whatever path is easiest for the physician and nurses.

Tolo distinguishes this from googling. She thinks the stigma around googling is partly deserved, because people have a cognitive bias toward catastrophe and fixate on rare outcomes like cancer. In her experience, ChatGPT gives more reasonable, better-calibrated answers: probably this, and in rare cases this. Mallin adds a caveat. The quality of the output depends on the quality of the input. He knows which questions to ask and which positives and negatives to include, and he isn't sure someone without medical knowledge would get equally good answers. He notes that ChatGPT has outperformed physicians on standardized tests, so it should do well in "standard" situations, but "life is rarely standard."

Tolo describes what she actually did. With her mother on the phone, she told ChatGPT to treat her as a patient and conduct an intake. It asked about symptoms, and she relayed perhaps a hundred questions back and forth. It then recommended exams and next steps. Her mother wasn't seen by an ophthalmologist in the ER, so she went to an optometrist for exams. According to Tolo, ChatGPT had essentially correctly identified the problem as some inflammation of the eye and a scratched cornea. Her mother is now on what Tolo believes are steroid eye drops.

The creatine study

The "topic of the week" is a recent UNSW study. Tolo summarizes it: taking 5 grams of creatine daily, a standard dose, did not produce more muscle gain than training without it. Over 12 weeks, both groups gained the same amount of lean muscle, about 2 kg. Early weight gain in the creatine group likely came from water retention rather than muscle. Johnson points out that creatine's audience is broader than gym-goers. He calls it one of the most commonly taken supplements, possibly the most common, and notes that it is used for cognitive benefits as well as muscle. The study measured only muscle.

Mallin finds the study's design interesting. The classic way to start creatine is a loading phase of about 20 grams a day for five to seven days to saturate the muscles. Because creatine increases fluid in the body, lean mass readings go up during that period. The researchers tried to account for this by measuring lean body mass at day zero and day seven, which they called a "wash-in" period, and only then began the 12-week comparison, with both groups following the same training plan. Mallin's reading is that earlier studies may have found muscle gains because they didn't separate out this water effect.

He raises several caveats. First, he thinks the wash-in was too short. At 5 grams a day without a loading phase, which is what this study used, saturating the muscles takes three to four weeks, not seven days. Second, he says the evidence that creatine increases muscle mass "was never actually that good," since earlier studies showed only small changes. In his view, the stronger evidence for creatine concerns cognitive function, bone, muscle-related longevity, metabolic health, and performance measures like speed and power. Those performance effects appear mostly in already well-trained athletes, not the untrained participants in this study. So what the study shows, according to Mallin, is that untrained people on a valid 12-week program won't see significantly different muscle mass with or without creatine. It says nothing about cognition, recovery, or power and strength output. He adds that 12 weeks is short, and that trained athletes followed for longer might show a difference.

Tolo says that, as someone "on the lighter side of science," she feels like a new study debunks the old consensus every week, which can paralyze ordinary people. Mallin's recommendation is to change nothing. By his account, creatine has more data behind it than almost any other supplement. The data is overwhelmingly positive, some studies show no benefit, and none show harm. He attributes the attention this study received to a culture that finds negative findings more interesting and likes headlines about accepted beliefs being overturned. Later in the episode he also notes that the study had about 60 participants, "so tiny," set against decades of research, and that new findings tend to be assumed better simply because they are new.

How much creatine to take

Johnson notes that Blueprint's Longevity Mix contains 2.5 grams a day, while many people take much more, especially during intense resistance training. He asks how to choose a dose and measure whether it works. Mallin says efficacy is hard to measure. For general health he would dose by weight at about 0.1 g per kilogram per day, which works out to roughly 7 grams for a 70 kg person. A flat 5 grams isn't precise, he explains, because smaller people, often women, may need less and people with more lean mass may need more.

Diet matters too. Mallin estimates that a generally healthy standard American diet already supplies 1 to 2 grams a day, so such a person might need only about 2.5 additional grams. Vegans and people who eat little red meat get less and probably need a higher supplemental dose, since the goal, at least for performance, is muscle saturation. If he had to give a single number, he would say 5 grams a day is adequate for an average-sized person with a normal diet. A loading phase isn't necessary, because saturation happens naturally over three to four weeks.

For cognitive purposes the doses are different. Mallin says there is decent data that creatine improves cognitive performance under high metabolic demand, such as brain trauma, poor sleep, or mild cognitive impairment and early dementia. Those studies used around 10 to 20 grams a day. He cites a study of up to 30 grams a day for up to five years with no significant downsides. As far as he can tell, excess creatine is simply excreted.

Tolo started with 5 grams when she joined Blueprint. Now she takes the Longevity Mix's 2.5 grams plus an extra half scoop to reach 5. At around 57 kg, she concludes that this is about right for her under Mallin's formula. Asked whether blood work can guide dosing, Mallin says not really. Creatinine, a kidney-function marker, will likely rise and appear falsely elevated with supplementation, but he wouldn't use it to judge dosing because hydration and kidney function can confound it. When Johnson asks how Mallin's own protocol will change because of the study, he answers, "Absolutely nothing."

Johnson says he weighs 77 kg and currently takes 5 grams: 2.5 from the Longevity Mix, and he is experimenting with 10. Because he usually has to get up around 3 a.m. to catch flights and sleeps less on travel days, he is trying around 20 grams on those days to see whether he feels more cognitively fresh. He has done this only twice and hasn't noticed a difference yet.

Tolo proposes something like a "travel advisory" for health. When a study goes viral, a central source would explain what it means, for example that a new creatine study shouldn't change your behavior, and why. Johnson agrees. His concern is that such headlines settle into the background of people's health consciousness as shorthand ("I remember it didn't really help muscle") and eventually harden into accepted truth.

Body awareness, and its strange side effects

Johnson describes himself as intensely body-aware after several years on his protocol. He says he can estimate his heart rate at any moment with pretty good accuracy, and that the many treatments he undergoes have taught him to inspect his body for color, tone, and function. He contrasts this with how unaware he used to be, when he would simply push through a headache. His explanation is that pairing measurement with how you feel builds fine-grained intuitions, effectively turning you into your own sensor.

Mallin raises a downside. He and his wife recently walked through a Las Vegas casino on the way to a show and wondered how everyone was still upright while drinking, smoking, and eating pizza and corn dogs. If he did that, he says, he would "literally be on the ground." Once you get healthy, he says, you can't allow yourself to slip, yet these people appear to be thriving without obvious back pain. He wouldn't choose differently, but he calls the position "unique." The response in the conversation is that the body is incredibly adaptive, able to postpone big problems for a long time before they come crashing down.

This recalls a study at Kernel, Johnson's brain-measurement company, on inebriation. Johnson explains that participants were tested sober and at low, medium, and high levels of intoxication. At low and medium levels, people could behave and pass tests as if they weren't intoxicated, but the brain scans showed impairment that the brain was compensating for. At high levels the compensation failed and impairment appeared in both behavior and the scans. Johnson draws two lessons. Cognitive decline begins long before symptoms appear, which makes early detection possible. And brain measurement can reveal things that self-perception misses. Mallin summarizes the analogy: you're "more drunk than you are aware," and the casino patrons may be closer to trouble than they feel.

Johnson then describes a recent team dinner of about 25 people at a well-known health restaurant in Venice, California. Watching the food arrive at tables around the room, he says he was "beside myself," given what he has learned about food, toxins, and how food moves through restaurant systems. His team teased him with "welcome to normal society." He says it felt like traveling back in time to watch a movie of the early 21st century, because the system they have built is so far outside that norm.

When will we know the protocol works?

Tolo asks how long it will take to know whether Johnson's protocol truly works, suggesting aging may behave like a cliff around age 70. Mallin says the cliff's location differs for everyone depending on genetics and lifestyle. For now, the best available approach is tracking Johnson's biological age and his speed of aging. Mallin reports that they have slowed his speed of aging to below 0.5, which he calls "pretty phenomenal." But to truly "not die," he says, the process has to keep slowing further. More still needs to happen to avoid a cliff at some point.

Karpathy's sleep tracker experiment

Tolo brings up a post by Andrej Karpathy, who ran a two-month n-of-1 experiment comparing four sleep trackers: Oura, Whoop, Eight Sleep, and Apple. He rated Oura and Whoop top tier. What interested Tolo most was his conclusion, which she reads aloud. Karpathy wrote that he could say "with absolute certainty" that Bryan is basically right and that his sleep scores correlate strongly with the quality of his work. With low scores he lacks agency, courage, and creativity. With high scores he can work 14 hours and barely notice time passing. He added that the effect depends on accumulated sleep over the last few days rather than a single night: one bad night is usually fine, but several in a row is bad news. Tolo says this matches her experience. People who sleep badly regularly can't tell what sustained good sleep feels like.

Mallin frames the post as evidence for measurement. Objective data helps you interpret subjective signals, and combining the two builds the knowledge to make good health decisions. Tolo's standard advice is to start measuring even without any interventions, so that you build a relationship between how you feel and what the data shows, and only then begin changing things. She jokes that Johnson should challenge Karpathy to beat his eight-month sleep score.

Johnson describes Karpathy as a founding member of OpenAI and former director of AI for Tesla's autonomous driving program, and calls him "one of the most formidable intellects of our time." He connects the post to his own thinking after selling Braintree Venmo. He concluded that humanity is at a pivotal moment, evolving into a new species, and that the most important thing to do was to improve our own intelligence so we could see the moment clearly and act wisely. That was the motivation for Kernel: measure the brain, reveal what is invisible, and pair it with AI to improve intelligence faster than otherwise possible. By intelligence he means not only IQ but emotional development, correcting blind spots, and avoiding narrow, tribal worldviews. As Johnson sees it, Karpathy is building the future of intelligence through AI and, through this sleep experiment, is also building his own. That an "architect of superintelligence" would endorse investing in one's own intellect through sleep and self-care felt to him like the best outcome he could ask for, and a reward for daily efforts urging people to stop eating junk food and drinking alcohol, stop going to bed late, prioritize sleep, exercise, and measure themselves. He says he greatly appreciated Karpathy's kind words.

Open questions about the format

The hosts end unsure whether the format works and ask listeners to comment on what they would change. Johnson notes they had six or seven more topics outlined that they didn't reach, including an embryo selection technology currently on the market. Mallin says he measures a podcast's quality by whether it was fun, and by that standard this one was.