Who Owns Code Security? Johannes Dahse on Developer Responsibility, Tooling, and AI
The Pragmatic EngineerWhat should software engineers know about writing secure code, and who in an organization should be responsible for it? In this episode of The Pragmatic Engineer, the host talks with Johannes Dahse, VP of Code Security at Sonar, who has worked in security for about 20 years. Dahse's central argument is that code security belongs to developers, not security teams. Vulnerabilities live in code, and only developers write and change that code. The conversation covers what that ownership involves in practice, which tools support it, how dependencies and developer machines bring in risk, and how AI coding assistants are changing both the threats and the defenses.
From Getting Hacked to Professional Penetration Testing
Dahse traces his interest in security back roughly 20 years, to when his own computer was infected by a widespread piece of malware. The experience was frustrating, but it also left him curious about how someone had gotten into his machine. During his school years he started experimenting with things like Trojan horses. Later he moved to a German university where IT security could be studied as a subject.
There he found capture-the-flag (CTF) competitions. University teams connect to an isolated online environment and try to hack each other for points. Dahse says he became "obsessed" with these competitions and calls them his best learning experience. Earning money from the same skills as a student followed naturally. He worked for a couple of years as a freelance penetration tester for large companies, then wrote tools to help with vulnerability hunting. That work turned into a startup, which Sonar later acquired.
He describes penetration testing as a simulated attack. A company hires a hacker to find vulnerabilities within an agreed time frame and scope. In his engagements the goal was usually to get into the client's network by exploiting weaknesses in the applications they ran. The job then was to document how an attacker could get in, not to go further or damage anything, so the client could fix it.
How a Penetration Test Actually Works
Dahse separates black-box testing, where the tester has no access or inside knowledge and behaves like a real attacker, from white-box testing. In a typical black-box engagement against a web application, the tester uses the application from the outside and tries to picture the code behind it. What is this code probably doing, based on what I can see? What might a developer have forgotten? Experience teaches testers the typical mistakes. The aim is always to reach something sensitive to the business, such as stealing files or getting into the database.
The host asks about tools such as port scanners. Dahse says automation mostly handles the mapping phase: finding which endpoints and ports exist. Once there is a good picture of the landscape, he used to work manually, poking at the application to "see what breaks when you touch it." He calls that the most exciting part of the work.
Developers Should Own Code Security
The host asks who inside a company should own code security. Dahse says the industry currently treats it as shared, and the dominant view is that the security team owns security because it is in the team's name. He sees it the other way around. Every software vulnerability shows up in code, developers are the only people in an organization who write and change that code, and they are the only ones who can fix security issues. So developers should own their share of the code and the security problems in it. He adds that this is more realistic today than it used to be, because developers now have good education and good tools available to them.
He does not think security teams are useless. Application security is a much broader field than code security. It includes compliance requirements, organization-wide security initiatives, vulnerability reports from penetration tests, and new emerging threats. The larger an organization gets, the more it needs a security team for this work. What Dahse objects to is security teams spending their time on every issue that comes up during development. He calls it a waste for a security expert to look at yet another new cross-site scripting issue, build an exploit and a risk assessment for it, when a developer could simply fix it while coding and move on. Taking that work off security teams frees them for problems where their expertise really matters, such as cryptography or authentication logic. In those areas they can also help developers.
The host compares this to the split between product teams and platform teams. A platform team has specialized expertise and builds things other engineers use. Engineers can come to them with questions like how to store two petabytes of data. Dahse agrees with the comparison, with one caveat: most of the ownership should stay with developers. Security should be part of the development process, not something bolted on or run ad hoc whenever the security team decides.
Why the Shift Happened, and Why Security Tools Had to Change
The host points out that developers historically were not expected to own security. Dahse agrees that security teams clearly owned it 20 years ago. The work was driven by compliance, and development cycles were much slower. A team might ship quarterly, with a security team doing a final audit just before release. Today teams release several times a day or even per hour, and AI coding assistants speed things up further. A disconnected review after the fact no longer fits that pace.
He says the tools have to change too. Security products were historically built for security teams, and they carry a matching philosophy. An auditor wants to know about every potential issue and to "turn every stone." Put that philosophy into a fast development workflow and it becomes too noisy, because developers get interrupted constantly. His analogy is a security expert sitting in your passenger seat and yelling about every possible danger while you drive. That might be interesting for the first 50 meters, then becomes painful and annoying. His conclusion is that developers should own code security issues, and the tools for code security should also be built for developers. Broader application security tooling can stay with security teams.
What Counts as Code Security
Dahse admits he lacks a better definition, but offers this: secure code is code free of anything an attacker can use to exploit the application, get to data, and put the business at risk. The hard part is deciding what counts as a "security issue." People usually think of classic vulnerabilities such as SQL injection. He argues the category is much wider. A null pointer exception that crashes an application leaves it in an unintended state, and attackers can abuse that in some scenarios. A more obvious example is memory corruption in C/C++, where a buffer overflow can let an attacker execute code on the server. There are also logic problems. If users can upload a profile picture, a developer has to make sure an attacker cannot upload a shell to the server instead.
His conclusion is that security issues are, in the end, just bugs: things a developer forgot, or things that were specified wrongly. That makes them a form of technical debt, not very different from other bugs in the backlog that a developer needs to fix. He says this framing also makes it clearer why security is a developer problem.
The Basics Every Developer Should Know
The host acknowledges that the topic is "mushy," ranging from obvious null pointer exceptions to less obvious buffer overflows, and asks for the basics. Dahse first narrows the scope. Developers need to know how to prevent and patch issues. They do not need to master full exploitation techniques or be able to run an entire buffer overflow attack chain.
His first basic is to really know and understand what your code does. He admits this sounds "a bit silly and obvious," but says it is exactly how security experts find issues: they hunt for corner cases and edge cases the developer overlooked. In an era of AI-accelerated development and heavy use of libraries and open source, he says, it is no longer a given that developers know what their code does and how it interacts with the rest of the codebase. His practical advice is to look at security-sensitive features through an attacker's eyes and ask what an attacker could modify.
Input handling is the classic example: never trust any input. Dahse notes that input can be subtle. When someone uploads a video to YouTube and sets its title, that title is attacker-controllable input. Developers need to track where GET and POST parameters, cookies, and other external input end up. Input that reaches a file operation might let an attacker open arbitrary files. In a SQL query it becomes SQL injection. In an HTML response it becomes cross-site scripting. These issues have been around for a long time, they are still the most critical ones, and they still show up.
His second basic is secret leaks, which he says play a role in many well-known data breaches. A developer hardcodes an API token, a cryptographic key, or a database password, sometimes only temporarily for testing. Attackers now crawl public GitHub repositories, collect secrets, and test whether they still work. Deleting the code does not help, because the secret stays in the git history. He says this still happens "because we are humans."
Asked for a checklist, Dahse says the list changes over time as the industry learns and development practices shift. Some issue types become less common and new ones become more common. The basics he described, though, have been around for a long time and "apparently they don't go away." The host later suggests keeping an eye on the OWASP Top 10 to cover the fundamentals, and Dahse agrees.
Advanced Issues and the Dependency Problem
More advanced issues usually need specialist expertise. Examples include encryption that an attacker can still break, authentication logic, access privileges, and password reset flows, which he says often go wrong. His advice for these complex security features is not to reinvent the wheel. Use solid, community-vetted frameworks and libraries, and ask the security team for help choosing them.
The host brings up package poisoning in the Node ecosystem. An attacker takes over a package, injects malicious code, and everyone who depends on it, directly or downstream, is affected. Dahse calls this "a tough one." Everyone uses dependencies, those dependencies use their own dependencies, and when a maintainer is compromised and a package is backdoored, a developer who pulls it in has almost no chance of avoiding the problem. Not using dependencies isn't an option. What teams can do is put tools in place, specifically software composition analysis (SCA). These tools are updated very frequently with known threats. Once a package becomes known as vulnerable, malicious, or backdoored, SCA warns you and tells you which version to move to, or that you should drop the dependency altogether.
He explains how SCA works. It reads manifest files, meaning the list of dependencies for whatever package manager a project uses, and checks them against a database of known problems. He distinguishes these known vulnerabilities (CVEs), found and reported in someone else's code, from the zero-day vulnerabilities a developer writes into their own code. As an example, SCA can flag that a particular Log4j version in your project is vulnerable to the known Log4Shell vulnerability.
The CVE Program and Whether Developers Should Track It
Dahse describes the CVE list as a database run by MITRE with US government involvement, and notes that "some change is happening here." It used to be the central database for documenting known vulnerabilities. He believes so many vulnerabilities are now reported daily that it has become a bottleneck, and other databases and collection points are appearing. SCA tools typically draw on the CVE database plus other sources to cover as many known threats as possible.
The host asks whether a developer at a scale-up with one security engineer should try to keep up with CVEs personally, or push for more dedicated security staff. Dahse's answer is neither: use a tool. This problem can be automated, and he would not hire more security staff for it. By his figures, the database holds over 200,000 CVEs and about 50 new ones appear each day, though not all of them are in open-source libraries. Some affect commercial products. Reading every CVE is a poor use of time for developers and security team members alike. The more important thing is that the tool helps with fixing, not just detection. Building a huge backlog of security issues matters less than being able to fix them and getting advice on how.
Findings from Sonar's State of Code Security Report
Dahse says Sonar's analyzers scan about 750 billion lines of code daily. For the report, the team studied a subset: 8 billion lines of code written by 1 million developers across 40,000 organizations worldwide. The headline finding was roughly one security issue per 1,000 lines of code. He says this matches his own sense from manual code audits, and he considers it "quite a lot."
The most common issue types were the basics already discussed. The top five included log injection, cross-site scripting, SQL injection, and hardcoded passwords. One surprise was how often regular expressions appeared: slow or insecure regexes that can enable denial-of-service attacks.
The host notes the irony of using lines of code as a metric, since its value is so often debated. Dahse says it is a statistic built for the report, but the underlying point is that code quality is tightly linked to security. The same problem can be solved in more or fewer lines. More lines mean more code to review and more places where security issues can hide, while well-maintained and well-structured code makes issues easier to spot.
Code Quality Is a Security Issue
Dahse calls the link between quality and security "totally underrated" in the industry. Some cases are obvious bugs, such as the null pointer exceptions and slow regexes he mentioned. The less obvious case is unreadable, poorly maintained spaghetti code. Code that is hard to understand is hard to review, so during pair programming or code review, peers are more likely to miss security problems in it. Fixing is affected too. When an issue is found and reported back, a developer working in unmaintainable code may struggle to fix it, and the attacker's window stays open longer. He sees this as especially relevant now, because AI-generated code, in his observation, is typically of poor quality.
He places code security within the wider field of cybersecurity: data security, cloud security, network security, forensics. Large organizations need all of them as interconnected lines of defense. From an offensive perspective, he has always found application security the most interesting, because every organization ships software or runs online services, and those are available to attackers 24/7. Applications sit at the front line and are typically the first entry point into a network. Other disciplines focus more on stopping what happens next, for example making stolen data impossible to decrypt, or preventing lateral movement. He explains lateral movement as follows: after an attacker gets an initial foothold, such as a shell on one server with the ability to run system commands, they probe the internal network for other services they can reach. A broader security strategy is needed to stop that spread.
Developer Machines, Supply Chains, and AI Agents
The host suggests that MCP servers could become a tempting attack vector. An attacker might publish an MCP server that claims to do one thing but secretly does another, runs locally, and gives the attacker access to a developer's machine. He asks whether developer machines have been off limits so far.
Dahse says they have never been off limits. Supply chain attacks are a major topic precisely because developers build software that gets deployed across organizations worldwide. Compromising a developer's machine is how a popular npm package gets compromised. A software vendor needs to make sure the software it ships to thousands of organizations is not backdoored because one of its developers was compromised.
He agrees that agents, and the greater control developers hand over to them, create a new threat. The concern is no longer just dependencies or general machine security. It is also whether the agent is doing the right thing, and whether it has privileges that would let it do harm, accidentally or deliberately. His example: an agent reads a Jira ticket, and someone has written a malicious ticket instructing the agent to add a backdoor instead of solving the development problem. That, he says, is a new type of security problem teams have to think about.
The Code Security Toolbox
Dahse walks through the tools he sees teams use for basic security hygiene. First is linting in the IDE, which catches issues as you type. IDE checks and extensions usually stay syntactic and semantic, and usually only cover the current file, because they must run in milliseconds without slowing the developer down. Security coverage there is therefore limited in breadth and depth.
Static application security testing (SAST) goes deeper. Depending on the tool, it may use techniques such as symbolic execution or taint analysis. The whole codebase is turned into an abstract model and the analyzer simulates what could happen at runtime without executing anything. He describes the model as a large graph in which every file, function, function call, and if/else, meaning every point where control flow changes, becomes part of the structure. The analysis looks for places where user input enters the application, follows it through variable assignments, branches, and function calls, and checks whether it reaches something security-sensitive. These data-flow paths can be very long and complex. He says the analysis used to take days and now takes minutes, and that making it efficient is a very hard problem. The value is that it automates the "be mindful of user input" advice from earlier and can find long, tricky connections between input and sensitive operations.
Other categories include secret detection for hardcoded credentials, and infrastructure-as-code scanning, which covers things like GitHub Actions files. He notes that a good SAST tool typically handles these too, since all of it can be treated as code. Then there is SCA for known vulnerabilities in dependencies.
The host asks about the downside of using every tool on every codebase, even for a one-person startup. Dahse recommends static analysis and SCA at every level as basic hygiene. The caution is to pick tools designed for developers rather than security teams. A noise level that is fine for an auditor is, in his words, "deadly" for development productivity.
Dynamic Testing: DAST and Fuzzing
Dahse also covers dynamic tools. Dynamic application security testing (DAST) tries to automate a penetration test. It treats a running application, on a test server or in production, as a black box and sends all kinds of malicious payloads at it. It then watches how the application responds: whether it breaks, slows down, behaves oddly, or throws error messages. Fuzzing is similar but aimed more at embedded software and C/C++ binaries and libraries that process complex file formats or protocols. It flips every bit of the input to see what crashes, and he says this works very well.
He considers dynamic tools less suited to developers today. They are disconnected from the act of coding: you have to finish your work, deploy it to a test server, and run the tool, so the feedback loop is longer and requires a context switch. For security teams, though, he calls DAST and fuzzing great additional tools. The host adds that the setup effort alone makes them a better fit for security teams.
The Limits of AI Security Reviews
The host asks about the many AI security review tools now on the market. Dahse finds them "super fascinating" and says what AI can find today is impressive. He then raises three limitations for systematic use by developers.
The first is precision and scale. Finding issues is only part of the job. It also matters how often a tool reports things that are not true or not meaningful issues, and whether it can handle something like half a million lines of code. What he mostly sees today are security research agents that find issues somewhat at random. That is useful for security teams, but developers need something that systematically finds all code security problems.
The second is determinism. AI is non-deterministic. That matters little to a security team, but a development organization needs a quality gate that produces the same output consistently across all teams. A build should not fail because an issue randomly appears or disappears.
The third is what he calls a contradiction. If much of today's code is written by AI, then using AI to review it is like having students grade their own homework. If the AI could not avoid generating a security issue, he asks, why would it detect that issue afterward? He argues for guardrails and verification that are not themselves AI. The host notes some people would answer that a different LLM could do the review, but agrees that this does not change the core problem.
AI-Generated Code: More Code, More Risk
Dahse describes Sonar studies of popular LLMs, including Claude, GPT models, Llama, and OpenCoder. The studies characterized each model's "personality": what kinds of issues it produces and what quality of code it writes. One finding he highlights is that GPT-5's reasoning mode decreased the number of security issues, though it did not eliminate them, but it produced more verbose output and more code to solve the same problem. He sees that extra low-quality code as its own security risk. Even if a snippet contains fewer issues, it may cause problems when combined with other code, it is harder for peers to review, and it is less maintainable.
The host links this to the old saying that code is a liability. Experienced engineers would sometimes spend a day or two cutting lines of code, refactoring, removing duplication, and clarifying responsibilities. Taking the one-issue-per-thousand-lines figure at face value, the host argues, that effort is still worth it. Dahse agrees that developers should not just "vibe code" and accept everything, but make sure the code makes sense, is well structured, and keeps the architecture maintainable. This already matters for quality, and it adds to the security problem. He cites a Stack Overflow survey in which, as he recalls, only 3% of developers said they trust their AI-generated code, and calls that very reasonable. The host shares their own habit when building APIs with AI: when a file such as an index.ts becomes bloated, they stop the AI and have it refactor, because they need to understand and navigate their code. They speculate that people who vibe code may eventually learn the same lessons the hard way.
How AI Is Changing Security
Asked about AI's overall impact, Dahse starts with security tools themselves. Deterministic algorithms like static analysis can be enhanced with AI. Taint analysis, for example, needs detailed knowledge of millions of libraries and frameworks, and AI can help gather that knowledge and feed it into the deterministic engine. He expects such combinations in static analysis, dynamic analysis, and other areas. AI also works well for fixing. Dropping half a million lines of code into a context window does not work well, but a narrow task does. If you give the model a deterministically found issue, the 20 relevant lines of code, and a description of the problem, AI is very good at fixing it. He stresses this matters because the goal is fixing, not only detecting.
He also sees applications themselves changing. In a traditional stack, removing the database removes SQL injection, but adding an LLM to the backend adds prompt injection, where attackers manipulate the system prompt or prompt logic and interfere with the output. As text becomes code, he says, "prompt injection is kind of like the new code injection," because human language is now the new code. Attackers adapt to the changing threat landscape, and the industry adapts too, though that can take time. He points to how much COBOL code is still around.
For coding assistants, he names verification as the big new challenge. Writing code is no longer the bottleneck. Verifying that all the faster-produced code is secure, at scale and at speed, is. Unverified code leads to security issues directly, or to quality issues that turn into security problems later. The host mentions the common answer of scaling human code review with better tooling and context. Dahse agrees that automated verification during code production is key. He also mentions work at Sonar on the training side: making sure the data LLMs are trained on is free of common security and quality issues, so that models produce more secure code from the start. He expects to see more of this in the future.
On whether AI introduces new kinds of flaws, Dahse says the models make the same mistakes as humans, since they learn from human code, but the prevalence of different issue types shifts. Slopsquatting is his example. An AI suggests a library that does not exist, an attacker registers that name on npm or Maven Central, and the malicious package ends up in the project. The underlying issue resembles the older dependency confusion problem, but a human rarely mistypes a dependency name, while AI makes the mistake much more often. Going the other way, he speculates, clearly labeled as something he is not yet seeing, that simple one-line issues like hardcoded passwords might decline as AI learns to avoid them. Meanwhile more complex issues, which only emerge when multiple code snippets combine, might grow, since AI finds those harder to grasp.
Misconceptions About the Security Industry
Having moved from the security side toward the developer side over the years, Dahse finds it striking that security and development are such separate communities. Both talk about code and bugs, but one is oriented toward building and the other toward attacking and breaking. He sees three fallacies coming out of that split.
The first is treating security as a product. The industry is fascinated by security problems, driven by compliance and fear of breaches, and full of money. It sells products that promise security. He says a CISO can have a hard time knowing what to buy. The mistake is choosing a tool that finds another 1,000 issues when you press the scan button, instead of something built into the development process that engages developers and helps them fix things. He calls this the biggest fallacy. The second is the mystique that security can only be owned by top-notch hackers. He argues the lines are blurrier than that, and the work is mostly about fixing, not the exploitation stage the industry talks about and finds fascinating. He admits he is guilty of that fascination himself. The third is that perfect security does not exist.
The host compares vendor promises in security to developer productivity tools that promise results from measuring one metric. In both fields, a team with no tools but experienced engineers can outperform a team with every scanner. The host wonders whether some areas are simply hard because they have many moving parts. Dahse agrees that the more complex the software, the more hard problems, bugs, and security issues it will have. He describes Sonar's vulnerability research team picking some of the world's most popular open-source projects, which have strong communities, good maintainers, and paid bug bounty programs, and still finding something every time. With enough motivation and a hard enough look, he believes you can always find something.
When Is Security "Good Enough"?
Given that perfect security is not possible, the host asks how engineers should decide when they have done enough. Dahse compares it to securing a house. Closing windows and doors will not stop a highly skilled, well-funded attacker, but it is basic hygiene you should have. The difference with software is that you add new windows and doors every day as you ship features, so the basics require automation.
His suggested process has a few steps. Start with an initial assessment of where you stand today, using professionals or a tool, and fix the most critical issues. More importantly, make sure new code does not add more vulnerabilities, or more technical debt and quality problems that later become security issues. Automation is key to this. After a quarter or so, run the assessment again, ideally finding that you shipped features without slowing down and that your security posture improved. He calls it "a never ending story." New developments like LLMs and prompt injection mean developers building on LLMs or calling their APIs have to ask new questions, "and the next thing will come."
In a closing question, the host asks which programming language he considers most secure. Dahse says newer languages tend to be more secure because they learned from the mistakes of older ones, and names Go as a good example. He adds that older languages keep evolving, and he considers Java, which is common in enterprises, quite secure to use.
What are code security basics that every software developer should know? Really know and understand what your code is doing. Maybe that sounds a bit silly and obvious, but that's how security experts find basically security issues in your code.
I'll set up an MCP server that says it does something, but secretly it does something else. It runs locally. Boom. With agents and giving away more control, there is a new threat here, because it's not just about the dependencies you're using or your machine security in general, but also making sure that the agent you're using is doing the right thing.
How do you think AI is changing code security and also security in general? Today, what we are seeing is...
As software engineers, what should we know about writing secure code? To answer this question, I turn to Johannes Dahse, who has been a security expert for 20 years and is currently the VP of Code Security at Sonar. In today's episode, we cover code security basics all software engineers should know of, common code security tools worth knowing of and using, like static application security testing, and more advanced tools like software composition analysis, how AI coding assistants introduce new risks and what we can do about these, and more. If you're a software engineer looking for pointers on how to make your code more secure, this episode is for you.
This podcast episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes to learn more about them and our other season sponsor.
So, Johannes, welcome to the podcast.
Thank you. A big pleasure to be here.
So, we're going to talk about cybersecurity today. I wanted to get to know how did you get into cybersecurity, and when was this?
It must have been like 20 years ago. I remember I got hacked, basically. My computer got infected. I think it was the Sasser one back in the days, and I was super frustrated, right? And then also super intrigued, like, how could someone get access to my computer? And so that led me into, you know, playing with security things like Trojan horses and that stuff at school time. And then I moved to Bochum in Germany, where you could study IT security, and that was exciting. And so they played capture the flag competitions, right? Those are hacking competitions where university teams connect online in an isolated environment and they try to hack each other to get points. And I got really obsessed playing those competitions, and that was the best learning experience for me, you know. And this led then into me getting into professional penetration testing, writing tools to assist the vulnerability hunting, and this led then into a startup which got acquired by Sonar, where I am today.
How can we imagine penetration testing?
So penetration testing is simulating an attack, basically, right? So you are kind of like... a company hires you as a hacker, basically, and you have to find out vulnerabilities in a given, you know, time scope and scope of the application that you should test. And it was just like a natural move, if you do that as a hobby with the hacking competitions, right, where you just do that for winning points in games, to then, you know, kind of like earn money with that as a student, to also look for security issues as a professional.
And I understand that, you know, like, yes, you can now hire penetration testers, right? As the company, I can hire teams that do this. But did you do some of this? Did you do professional penetration testing?
Yes, absolutely. For a couple of years, doing this as a freelancer and for big companies, basically, right, looking for security issues they have and always trying to get into their network, typically by exploiting vulnerabilities in software applications they're running, and then documenting this, right? Not going further, not destroying something, doing something malicious, but basically reporting then: this is how an attacker could get in, so they can fix it.
How does the penetration test look like from when the company says like, "All right, come and penetration test us"? How do you actually go around? Do they actually give you access to some of their systems, or do you need to just assume no knowledge? How does that work?
Yeah, there are different types, right? So, there's black-box penetration testing and white-box. Black-box meaning you don't have access to anything. You treat it as a real attacker, like with no knowledge. Typically you have maybe like a web application running, and you go and look at it from the attacker perspective and play around with the application from the outside, trying to imagine basically what could be the code behind that application, what could the code be doing that I'm seeing from using that application, and then trying to figure out what could be vulnerabilities here, right? What could be something a developer forgot to do? And by experience you learn a bit what are typical mistakes, and then you go from there, always trying to, you know, exploit something where you can steal files or get access to the database or something where there is some data, something, you know, that could be sensitive to the business, that is then security critical.
Do you bring your own tools there? Do you have like methodologies? Like, I guess, you know, it's very basic, but I do know of the concept called port scanning, where you write software that tries all the different ports. It sends messages, and you hope that if they configured a server incorrectly, or maybe correctly, you can get through. But what kind of tools do you come with? Do you use tools as a penetration tester?
I think mostly for kind of like mapping out what's available, right? That's, I think, the biggest part you automate, so you don't, you know, test for all the endpoints or ports, as you said. But I think then, once you've found a good landscape of what's out there, you go in manually. At least I used to do that manually, to, you know, try to poke and see what breaks when you touch it and when you play around with it. And that's also the most exciting part, I think.
You know, in the real world, we're going to be sitting, you know, as software engineers inside a company, and we're going to be building our software. This might be services, this might be apps, websites and so on. And there's going to be attackers outside. It's going to be like script kiddies, or, you know, like people just malicious, playing around, poking around, and there's going to be professionals as well who will be trying to get financial gain, whatnot. Now, inside the company, who should own code security?
You know, in the industry today, what we're seeing is that this is a shared responsibility, basically, right? So we talked about the penetration tests, and typically security teams are involved in that. And then there are still developers, you know, adding and writing code, right? And I think predominantly in the industry the view is that the security team should own all that security, right? It's in the name of the team. But I see it quite the other way around. I think, you know, every software vulnerability basically manifests in code, and developers are the only ones writing the code in organizations and changing the code, right? And they're the only ones who can fix security issues. And so I think they should own all those code security issues, right? They should own basically their share of the code and also the problems related to their code. And I think that's also more realistic today, because you have great education available and great tools available for developers. So I think that ownership should be with developers on the code security problems.
I hear what you're saying on devs should own code security, but then why have a security team, or at what point should you have a security team? Again, you now work with a lot of different companies and sizes, and previously you also worked as a security engineer. At what point do companies bring in a security team, and when they do, what is their role? I'm kind of like, look, if there's a security team, I'm like, come on, that's their name, right? Like, as a developer, security as a whole, like to make your service 100% secure, that's pretty daunting.
Yeah. I don't think that security teams are useless, right? Not at all. We talked about the penetration test; that's typically something run by security teams. And so I think the field of, you know, application security is just much broader than code security, right? So you have maybe compliance requirements that you need to look after, and some, you know, organization-wide security initiatives, or there are vulnerability reports coming in from a penetration test, or new threats are available. And so security teams, I think, should look at this broader application security field, and it's good to have a security team for that, and the larger the organization gets, right, I think you need a security team. I just think that when you write software and when organizations, you know, deploy software, the security team shouldn't waste their time looking into every single security issue that happens during development, right? And I think that part should be fully owned by developers, right? I think it's a waste of time to look at every single new cross-site scripting issue again and again and try to exploit it and build some fancy exploit and risk assessment, where developers could just, you know, fix issues as they code and move on. And that also allows security teams to have more time to actually focus on bigger problems, on problems where they can really bring in their expertise, like cryptography or authentication logic or things like this, where they can then also be very helpful with their expertise for developers.
So I'm kind of hearing some similarities between the kind of feature teams or program teams and platform teams, where, you know, platform teams typically build platforms that engineers can build on, and they have a specialized expertise. It might be a massive database platform, like for a large data storage company, and then engineers kind of use the APIs, but they don't need to know all the details. But when they do, they can just go to the platform team saying, hey, how do I store, you know, like two petabytes of data, and they'll be like, okay, here's different ways you can do it. So do I understand correctly that you're kind of saying security teams will also be this like specialized expertise, where they can help you with a bunch of stuff, and they will try to build tools as well for devs to like self-service, or, you know, share common things to watch out for?
Yeah, exactly. I think it's a good comparison, right? Definitely helping, but also leaving, you know, the majority of ownership there with developers, so they can basically have security as part of the process of development, and not just something that is, you know, attached to or ad hoc run whenever the security team decides, right? I think it should be really part of the process of development, and that must be then owned by developers to really, you know, engage in security issues and fix them, because that's what makes you secure in the...
I will challenge, though, that historically I don't think software engineers owned security or were expected to own it. Can we talk about how this changed over time, over your 20 years, how you've seen changes happen? Because I do feel it's shifting left onto developers, but what was the historic context here, and what is changing now?
Historically, I think it was clearly owned by security teams, right? So if you imagine 20 years back, it was all about compliance, driven by compliance, and then also the software development life cycle was a lot slower than today, right? You would have your quarterly release, and before that a security team would come in and do a final audit, right? And then you would release. And today we're just moving at a much faster pace, releasing a couple of times a day or per hour, right? And you have AI coding assistants, and so we are moving a lot faster now.
Johannes just talked about how engineering teams today are moving at a much faster pace than before, especially teams using AI coding assistants, which are most engineering teams, honestly. Here's something surprising, though. As dev teams build more products and features faster than before, coordination is increasingly the problem. You now have more Slack channels pop up, more customer feedback to deal with, and you often end up switching between different tools to decide what to build and how to build it. This is where our season sponsor, Linear, can help dev teams stay focused.
Sierra is an AI-powered customer experience startup. They were preparing for the next phase of company growth and wanted to find a tool that can help a larger team move quickly without slowing down. They chose Linear as the operating system of the company and wired all of their work into the platform. Today, project updates in Linear ripple through Slack, customer requests are logged in Linear, and stats from Linear are pulled out into company dashboards and into the slides that Sierra shows off as they celebrate wins at all-hands meetings. Despite Sierra being in hypergrowth, everyone understands what they're building, why they're building it, and how the work is progressing. What I love about Sierra's approach is how they didn't set up Linear wanting to know what individuals did in a given week. They wanted to know what was accomplished in service of which projects. This is the beauty of using Linear. It helps hypergrowth companies stay focused, spend more time building, and less time coordinating. If your team cares about tools that remove additional work for the team instead of adding extra to it, check out Linear at linear.app/pragmatic. And now let's get back to fast-moving engineering teams and security reviews.
You cannot have this disconnected security review that you do afterwards. And so what's also changing in the industry here is, I think, the tools that you need for this. You know, historically the tools were built only for security teams, right? And with that there is a different product philosophy that comes with security products, because as a security auditor, basically, you want to know about every single potential issue, right? You want to turn every stone, and better look twice than never, to find out, you know, what could go wrong. And now, if you apply this to this new pace of fast development, that doesn't work anymore, right? Because you can't get interrupted all the time with findings. It's too noisy, right? I like to compare this with, you know, if you drive a car and you have a security guy in your passenger seat, and he would scream and yell at you at every single thing that could go wrong all the time. That's maybe interesting the first 50 meters, but then gets super painful and annoying.
I think with that we see a change in the industry that, you know, I think developers should own code security issues, but also the tooling around code security issues must be owned by and built for developers. And then there are other application security tools, and application security as a broader thing, that should still be owned by security teams.
So, so far you've mentioned two different things, if I caught it correctly. One was code security, and then application security, and you said that application security is a lot more than code security. So it's, you know, like a superset of it. What is code security? I mean, this is one of your areas of expertise, as I understand, but how do you define that? You know, where does it start? Where does it end? Because it does sound like something that, as software engineers, we should be aware of, right?
Yeah. For a lack of a better definition, I would say it's basically code that is free of security issues, free of anything that can be leveraged by an attacker to exploit your application and then get access to some of your data and put your business at risk.
But with that simple definition, I think the complexity is a bit: what are security issues when we say code is free of security issues? And I think here we think typically of vulnerabilities, right? SQL injection is a vulnerability. And I think it's much more than this, right? If you think about bugs like, I don't know, a null pointer exception where your application crashes, then your application is in an unintended state and this can be abused by attackers in some scenarios. Or maybe a more obvious example would be memory corruption problems in C/C++, where as an attacker you can do a buffer overflow and then execute code on your server.
And so I think here the lines get more blurry. And then there are also more logical things, like if you write an application where you can upload a profile picture, you shouldn't forget that an attacker shouldn't be able to upload a shell to your server, and those kinds of things.
So I think we are realizing that code security is much more than just vulnerabilities, and in the end those are just bugs, right? Those are either things you forgot about in your code or those are misspecified things. And so it's basically technical debt, right? It's not so much different than other bugs in your code that you have in your backlog and you just need to fix as a developer. And I think from that perspective, it's also more clear why that's a developer problem and should be owned by developers.
I understand you. We should be owning code security, but it's a pretty mushy subject, as you say. It's a lot of things, from the obvious null pointer exceptions to maybe the not so obvious buffer overflows, which are a little bit harder to work with if you're not aware of it. Of course, sometimes you use languages that solve for it. As a software engineer, what are code security basics that every software developer should know, in your mind?
You just mentioned buffer overflows. I think the key here is, for developers, in those basics they need to only understand how to prevent those issues and how to patch them. They don't need to understand the full exploitation techniques to run a buffer overflow attack, right? Like, you can patch things without necessarily needing to run the full chain.
And I think of some of the basics you should be aware of, the first thing that comes to my mind is to really know and understand what your code is doing. And maybe that sounds a bit silly and obvious, but that's how security experts find security issues in your code. They try to look for corner cases and edge cases that you may have forgotten about or overlooked.
And maybe in the time of AI-accelerated development and using libraries and open source code, it's not so obvious anymore to say that we all the time know what our code is doing and how it interacts with our code base, right? So I think one thing we can do here is to really look through the eyes of an attacker, at least when working on security-sensitive features. What could an attacker do here, and how could an attacker modify something here, right?
The industry has been talking for a long time about this input validation, input sanitization, right? Maybe that's a good example here, where never, never trust the input, right? Any input. Yes, exactly. And this can be also a bit more subtle, right? Like if you upload a video to YouTube and someone parses with their application the YouTube video titles, then that's input basically, right? Because you modify the YouTube title name.
But then really making sure we think about this: where is all that, GET parameters, POST parameters, cookies, external input used? And where am I using this in my file operation, which could be modified to open arbitrary files by an attacker? Or traditionally, in a SQL query you have a SQL injection, in your HTML response page you have cross-site scripting, and those typical things. And I think we are still seeing those issues, right? They are the most critical ones and they have been around for a long time, but we still see those issues.
And then secret leaks, I think, is another basic thing that is involved in many popular data breaches, where a developer hardcoded, maybe just for testing purposes, temporarily added, like, hardcoded into the code a little API token.
So like the secrets, like API tokens, well, like all sorts of tokens, right? That should typically live in your local environment variables.
Exactly. Exactly. And it can be API access tokens or cryptography tokens or passwords for the database or whatever. Attackers nowadays crawl the public GitHub repositories, right? And steal those secrets and try to see if they're still valid. Even if you delete your code, right, it's in the git history and it gets parsed. So I think that's another basic thing we should be aware of and not do, and it still happens because we are humans, right.
So these were the kind of, I guess, the basics to cover as a developer. Is there like a checklist I could go through? Because again, you listed a bunch of them, and depending on your level you either say these are super basic or, like, what are these things. But the parameters, SQL injections, secret leaks and some other things. Do you have a go-to list of, like, go through all these things and make sure you understand each of these things and you can check your code or know if they would be applicable?
It changes a bit over time also. We are evolving and we are learning more about certain security issues, and certain types of issues we do less, and then maybe new types are becoming more prevalent. Maybe also because of how the landscape changes or how development changes. But again, I think those basic ones we talked about have been around for a long time and we still see them. Apparently they don't go away.
And what about the more advanced things that could go wrong? Because these were the basic ones, right? I think we just covered the basic ones, but you must have seen some more exotic security issues that maybe would have not been as easily preventable, or a lot more creative ones.
So there are more advanced things in terms of maybe the expertise that is needed. If we talk about cryptography things, right? If you're encrypting something and an attacker is still able to decrypt it. Or there is some authentication logic or access privileges or password reset functionality, which is also something where often things can go wrong. I think the key as a developer, for those more complex features, security features, is to not try to reinvent the wheel and just use solid frameworks or libraries, something that is vetted by the open source community and trusted. And I think here again a security team can help you with that, right?
One of the recent security issues that is coming up in the Node ecosystem is packages being poisoned, where an attacker takes over some packages, they inject malicious code, and whoever is using a package, or the downstream dependency of that package, they can be impacted. I think we've seen a crypto-related issue like this. In your view, who could best protect against these issues? Would it need to be a security team who decides on things like pinning certain versions of packages or scanning updates for it? Or basically, as a developer, if I'm depending on third-party packages, what are good practices I can do to try to avoid some of these dependency security issues, which are now becoming more widespread?
That's a tough one, right? Because everyone uses dependencies and your dependencies are using dependencies. And so it's quite hard to do something, right? If you have this whole dependency chain and some developer of that dependency, a maintainer, gets compromised and then a dependency gets backdoored, you have almost no chance in having a security problem when you pull in that dependency. You cannot not use dependencies.
I think the only thing you can do here is to have tools in place, and this is like software composition analysis, which is a thing here, that basically observe and check your dependencies for known threats, right? At some point, luckily, like the npm package you mentioned, it became known to be vulnerable or malicious or backdoored, and then those tools basically get updated on a very frequent basis to look at what are the threats and what are dependencies you shouldn't be using in a specific version, and then warn you about this, and what is the next version you should use, or that you should get rid of that dependency basically.
And what is software composition analysis?
So software composition analysis, called SCA, is basically a technique where we look at manifest files, your list of dependencies, right, depending on the package manager you use. And then this list of dependencies is checked against a database of known security problems, right? Those are called the CVEs. Those are not the zero-day vulnerabilities we talked about earlier that you typed into your code, right? Some maintainer had a security problem. Someone found that problem, reported it, it's documented in a database, and then you can basically, with software composition analysis, map that this specific Log4j version of your library is vulnerable to the Log4Shell vulnerability that is known, and then it can warn you.
And can you tell us about the CVE program? I understand inside security circles this is very well known and very useful, but what should I know as a developer about this, and how much should I kind of look it up, check it, worry about it?
It's run by MITRE, right, like the US government. There is some change happening here. So that's the common vulnerability enumeration, the CVE list, and it's a database where, it used to be kind of like the central database for documenting known vulnerabilities. I think it's just too many vulnerabilities reported every day, so I think there's a bit of a bottleneck there. And so there are also other databases evolving, or places evolving, where security issues are collected and gathered. And SCA tools typically use that CVE database, but also other resources, to collect all kinds of known vulnerabilities to make sure they know about all potential threats.
And as a software engineer strictly focusing on, you know, I'm trying to make my code secure, do you see value in trying to keep up with CVEs, with new vulnerabilities? Or do you see this being more something where you really need someone who is dedicated, focused on this, maybe a security engineer? I'm just talking from a practical perspective. If I'm working at a scaleup where we have a midsize team, maybe have one security engineer, and I really, really want to do my best work, security is important in our domain. Do I take some of this on me, or do I say, hey, if we really need this, let's get more resources, dedicated folks who can help with the kind of depth of the industry?
Yeah, I would use a tool for this, right? It's a problem that you can automate, and I wouldn't hire more security team members for this. So you can use software composition analysis. It will automatically check all the dependencies. There are, I think, in the database over 200,000 CVEs, and every day I think there are like 50 new CVEs coming out, not necessarily in open source libraries, right, also in known products, etc. But I think it's not a good use of your time as a developer, but also not as a security team member, to look at every single CVE that comes out. I think you should then have a good tool in place, a software composition analysis tool, that helps you to detect those but also helps you in fixing those, right? Which is much more important than building a huge backlog of security issues. The important thing is that you can also fix this and get some advice on how to fix this.
Johannes has just talked about how it's a no-brainer to automate much of your security analysis, like keeping up with the latest security vulnerabilities in software engineering. Using the right automation and the right tooling means that you get to focus on what matters, like building your product, and not spend as much time on infrastructure. This is where our presenting sponsor Statsig comes in.
Statsig gives engineering teams a toolkit for safer deployment: feature gates, gradual rollouts, and experimentation. These are built into your release process. So you ship changes to 10% of users, then expand to the remaining 90%. You validate behavior, measure real impact, and scale only when things look good. If something goes wrong, you can instantly turn it off before it affects everyone.
To support this, Statsig includes product and infra analytics, built-in tools for logging and tracing, so you can actually see what your code is doing in production: performance, errors, user behavior, all in one place, because you cannot secure what you cannot observe.
For teams with strict data governance or security requirements, Statsig also offers warehouse native. Your user-level data stays in your data warehouse, Snowflake, BigQuery, Databricks, whatever you use. Full control inside your security boundary, and you get the deployment safety and observability without shipping sensitive data to external systems. Companies like Microsoft, Atlassian and Brex use Statsig for safer deployments with enterprise-grade security. Statsig has a generous free tier to get started, and pro pricing for teams starts at $150 per month. To learn more and get a 30-day enterprise trial, go to statsig.com/pragmatic.
With this, let's get back to code security with Johannes. And recently you've produced a State of Code Security report, which is a pretty comprehensive one, as I understand. What are things that you found there?
Yeah, so at Sonar we scan 750 billion lines of code daily, right? So our analyzer sees quite a lot of code, and we studied a subset of this. We took 8 billion lines of code that was written by 1 million developers of 40,000 organizations globally. Quite a data set. And then we looked at what are the issues we see, and I think one finding was that about every 1,000 lines of code we see a security issue, and that reflects kind of well my feeling from when I manually audited code.
So an issue every 1,000 lines of code is quite a lot, I think. And then the issue types we found and saw were the basic ones we talked about, right? In the top five at least, there was log injection, cross-site scripting, SQL injection, hardcoded passwords, the typical things that go wrong. I think some surprises were in there. Regular expressions, for example, was something that apparently more often, you know, we have a slow regular expression or insecure regular expression, which can lead to denial of service attacks. And so that would be something more out of the lines. But yeah, the basic ones are still very prominent today in code.
It's very interesting because, like you're saying, every 1,000 lines of code, roughly one security issue. It's funny because lines of code we always argue about is
it a good measurement of things, you know, complexity, work, whatnot, or is it not? But I guess you're still using this heuristic, right?
I mean, it's a statistic we build for the report, right? But I think, yes, what it comes down to is that quality here is really connected to security, right? I mean, you could solve certain problems with more lines of code or with fewer lines of code, and I think quality is something here that is very related, in terms of when you have more lines of code, there's more code to review, basically, and it's harder to spot security issues in the end, while if you do it in a well-maintained and structured way
Exactly, this was exactly my feeling on this. So what would you say, how is the quality of code related to security? Did you see any findings on this?
Yeah, I think it's super related, right? And I think it's totally underrated in the industry today. We talked about the null pointer exceptions or these slow regular expressions, right, that can lead to security issues, and those are maybe the more obvious examples of bugs. But also, if you think about unreadable code, not well-maintained code, kind of like spaghetti code, then it's not so obvious at first maybe how this is connected to security. But then if you think about that code not being easy to comprehend, not easy to review, and you do pair programming or code reviews in your development team, then in that spaghetti code you will more likely oversee security problems of your peer. And then also if you think about fixing security issues, right, like maybe someone found an issue, or found an issue later, and reports that back to you, and you as a developer have to fix it. Think about if that's not well-maintainable code: you cannot fix the security problem. So quality suddenly becomes a security issue in the sense that the attacker window stays open longer. At some point you have to fix the issues. And so I think here code quality is super related to code security, especially now with AI-generated code, where we typically see poor quality of code, right, and that becomes a problem for security.
When we look at code security, how does that relate to cybersecurity as a whole?
So there are many fields of security, right? There's data security, cloud security, network security, forensics. As a larger organization, you kind of need all of them, and they are all interconnected and they build multiple lines of defense. From my perspective, from an offensive security perspective, I always found application security the most interesting field, because if you think about it, every organization today basically deploys software. They ship software as a product or they deploy some services online to have customers interact with their business. And so those applications are online 24/7, right? And they're available to me as an attacker. And that's at the forefront of security and typically the first entry point into the network. And so that makes the application security field so critical, or so interesting for attackers, whereas the other areas more try to prevent the lateral movement once an attacker is in. Can the attacker maybe not decrypt the data he stole, or can the attacker not move from one server to the other?
What is lateral movement?
Yeah. So typically as an attacker you would gain your first entry point into a network, and then maybe you want to expand from there. So you have a shell on one server. You can control a server or a machine. You can run system commands, and then from there you are in the internal network and try to see what other services can I reach, what other internal things can I access. And then you need a security strategy, basically, in the broader cybersecurity strategy, to prevent that lateral movement between internal services.
One idea that comes to me about lateral movement with the advent of AI assistants and MCP servers: it's probably going to be a pretty tempting attack vector. Just thinking as an attacker: hey, let me try to get access to that developer machine. I'll set up an MCP server that says it does something, but secretly it does something else. It runs locally. Boom, I get access to this developer machine. As developers and as security professionals, how much should we worry about this? And are you seeing any worries about this specific attack vector? Because I feel until now developers' machines were kind of a little bit off limits, or were they off limits?
Yeah, I mean, developers' machines I think are not off limits, right? I think supply chain attacks are a big topic, where developers are building software and then that software is deployed at organizations worldwide, right, and that makes it so interesting. So we talked about an npm package that gets compromised by compromising a developer's machine, basically, right? And then from there you can compromise a super popular dependency. Or if you're a software vendor, you'd better make sure the software that is shipped to maybe thousands of organizations is not backdoored because some developer got backdoored. And yes, I think also with agents and giving away more control, there is a new threat here, because it's not just about the dependencies you're using or your machine security in general, but also making sure that the agent you're using is doing the right thing and doesn't have the privileges to do something harmful, accidentally or on purpose, as you said. Like if the agent parses a Jira ticket, someone can create a malicious Jira ticket that basically instructs the agent to add a backdoor instead of just solving a development problem. Then you suddenly have a new type of security problem to think about.
You previously mentioned that if you can automate things for code security or application security, you should try to do that. What are the common code security tools that you keep seeing engineering teams use for security hygiene? What are the categories?
I think every developer uses an IDE, right? So there's some basic linting available in IDEs, and that's great, because as you type you find issues and you can resolve them. It's just that in an IDE you typically don't have such broad or in-depth security coverage built in. There are some IDE extensions you can use, but then typically you stay on the linting side, which means some syntactical and semantical checks, and typically in the current file you're working in, simply out of performance reasons, right, because it has to be done in milliseconds as you code and shouldn't slow you down. And then you have static application security testing tools, SAST tools, that can go to a deeper level of code analysis. Depending on the SAST tool you use, there are for example symbolic execution or taint analysis techniques used, where your whole codebase is transformed into an abstract model, basically, and then static analysis is simulating what could happen here at runtime. It's not executing the code, right, but analyzing it, and connecting to what we talked about earlier: user inputs, for example, how are they flowing in terms of data flows through all your code paths, and simulating what could go wrong here to find different issues.
And can you just give us a high level of what is happening? Because this sounds super interesting. What I understood, and tell me if I got it right, is you take your code and you kind of turn it into maybe a graph of some sort, and then you can try to figure out inputs, how they can flow, how they can get to components.
Yeah, exactly. So your code is transformed into a big graph model. This can be of any dimensions. Yes. So basically every file of your codebase, every function, every if-else, so whenever the control flow of your application changes, every function call, every if-else is part of that big graph model, right? And then you try to figure out what are all the combinations of your variable assignments, which create data flow, basically. Where is user input received in that application and then passed on with data assignments through different if-else and function calls, and where does it end up in something security-sensitive? And this can be a very, very long data flow path and very complicated to do, and also to do efficiently, right? It used to take days, and now we can do that in minutes, and that's a very hard problem to solve. But it helps you to automate that process, right? What we talked about earlier, where you should be mindful of what is user input: it helps you to automate that and find even very tricky and long connections between user input and something security-sensitive.
Okay. So we talked about the linters inside IDEs, the SAST scanners that you said. Are there other tools worth knowing about?
I mean, secret detection. We talked about hardcoded passwords, so there are secret detection tools. There is infrastructure-as-code scanning, right? If you think about code more broadly, it's also infrastructure as code, or your GitHub Actions file can be code, right? And there are tools to scan this. Typically if you have a good SAST tool, that's all covered by static analysis, basically, right, because everything can be considered code here. And then we already talked about software composition analysis as another tool for developers, where you find those known vulnerabilities, those CVEs, in your dependencies.
I guess this is a layered approach, right? So the more security you'd like, the more of these layers you would set up. But do I sense that there's a trade-off between them? It's going to be maybe complexity, time to run, those kinds of things. What is the downside of just throwing all of these tools onto every single codebase I have, even if I'm a one-person startup? Why would you not recommend that, if you would not recommend it?
Yeah, I think for the basic static analysis tools I would definitely recommend doing that. I think what you should be careful of here is choosing something that is intended to be used by developers and not by security teams, right? We talked about the noise level that is interesting for security teams from an oil perspective, but deadly for your product development productivity, where you shouldn't be annoyed. I think that's something to watch out for. And then there are differences for SAST tools and SCA tools, whether they are more for security teams or for developers. But I would definitely recommend, I think at all levels, to run static analysis and software composition analysis to have your basic security hygiene in place.
So these are static tools. After you write the code, they can run on it, they can run on CI, they can run with continuous deployment. Are there more dynamic tools? I'm just thinking of the idea that as your code runs, as your servers operate, they dynamically try to test or just do funky stuff.
Absolutely. So there's dynamic application security testing. We talked about penetration tests, right? And a DAST tool tries to automate exactly this. What is DAST? Dynamic application security testing is testing from the outside, as a black box, when your application is already running on a test server or in production. And it's basically shooting all kinds of malicious payloads from the outside against your application to see how it reacts: is it breaking, is there a delay, is it behaving weirdly or throwing an error message? And this way it tries to automate such a human penetration test to find out if there are issues it can detect. And then there's also, on the dynamic side, fuzzing, which is similar to DAST, basically, where it's more for embedded software, binaries, C/C++ libraries or applications, where typically you parse complex formats or protocols like file formats, and then you want to flip every single bit, basically, in what you're processing to see if something breaks, right? And you can automate that with fuzzing and then find those crashes. So that works very well. I just think that those more dynamic tools are not so much for developers today, because you are a bit disconnected from your coding and you have to context switch, basically, because you cannot find things as you type. You need to kind of finish what you're doing, deploy it on the test server, get it running, and then the feedback loop is just a bit longer. And so I think for developers it's more inefficient, but for security teams it's a great tool to have, to additionally maybe run a DAST or fuzzing, right.
Yeah. And as you say, it sounds like a bunch of setup to do just one more thing. I can see why you're saying that it's more for a security team. One thing you haven't mentioned, and I was waiting to see if you would mention it: AI security reviews. These are popping up everywhere. There are a lot of different tools, a lot of different vendors, some existing ones, and they're all saying the same thing: use this thing, it will make your code more secure. What is your take as a software professional?
I think it's super fascinating and fun to see, right? And also impressive what AI can find today. As with static analysis or every other technology, to me it's not necessarily all about just finding issues, at least when you want to use it in a systematic way as a developer. Here you have to get into a good balance of: am I not only finding things, but how often am I reporting things that are actually not a true or meaningful issue? And can I scale this to half a million lines of code, etc.? So I think what we're seeing more today is security research agents that go in and randomly find issues. It's a great tool for security teams here. But as a developer, you want to have, I think, something a bit more systematic that finds all code security problems. And to me there's another aspect of being deterministic versus non-deterministic, right? Here the AI basically is non-deterministic. And again, for a security team that's not so important, but as a development organization you basically need to have a quality gate that's consistent across your team and all the other teams, that always has the same output, and you cannot fail your gate because a new issue is randomly popping up or disappearing, etc. So I think that doesn't really work well for developers today. And lastly, to me, from a developer perspective, there's also a bit of a contradiction if you think about how most, or a lot of, code is written by AI itself, depending on who you ask. If you then use AI to review AI-generated code, that's a bit like having students grade their own homework, where you could think that if AI couldn't prevent actually generating a security issue, why would it then detect that security issue? So I think we need to have good guardrails and verification in place that is not AI, to then verify this AI-generated code, basically.
I can see where you're coming from. Although I can also see some people might say, well, what if it's a different AI? What if it's a different LLM? But we're still not, you know, we're just
chasing one another like we're still not changing the core problem here.
And earlier you said that AI can generate low-quality code and this could be an issue when we're talking about the lines of code per security issue. Can you go a little bit into more detail on what you're seeing, observing? What does low quality mean in this sense? Or is it the verbose nature of that we're sometimes seeing?
So for example, at Sonar we did great studies of the most popular LLMs like Claude, GPT-4, 5, Llama, OpenCoder, etc. And we looked at what we call personalities of them, right? What kind of issues do they produce and what kind of quality are they producing? And then measured what comes out of that and studied this. And one interesting finding to me was, for example, that if you use the reasoning mode of GPT-5, it actually decreases, not eliminates but decreases, the number of security issues you find, but it's using more verbose output to solve the development problem. It produces more code actually, right. And this is then again something that leads into security problems because you have more low-quality code that maybe has less security issues itself but then poses a problem maybe combined with other snippets of your code, or it's harder to review by your peers or later on, and it's less maintainable and leads to a security problem.
This reminds me of there's this old saying before AI that code is a liability. The more code you have, the more liability you have. And this was a reason that back in the day an experienced engineer would sometimes spend a day or two reducing the lines of code, refactoring, compressing it, bringing single responsibility, removing duplication. And I wonder if we've kind of forgotten this a little bit, that the more lines of code you have... I mean, just taking your statistic of one security issue per thousand lines of code, let's just take it for now, you would want to have kind of efficient lines of code, right? You do want to spend that time and effort of getting to a system that is simple, clear responsibilities, concise.
I think this is something that developers or engineers look at today already, not just to vibe code and accept all the code, but to actually make sure that the code makes sense and is well structured for all kinds of purposes, right? To maintain a good architecture in your codebase and to have good maintainable code, etc. Outside of the security world, that's already a big code quality problem, and I think developers are aware of this. But yes, it adds to the security problem. On top of that, there was a nice survey by, I think it was Stack Overflow, where I think only 3% of the developers asked basically said they trust their AI-generated code, and I think that's very reasonable.
Yeah. I mean, when I'm using AI to build myself my APIs and tests and all those things, I also find myself, I give it instructions and every now and then I also tell it to refactor some things, to move things around. As I'm watching the output, I just see something is getting bloated. You know, I have an index.ts that is getting this big. I'm like, "All right, let's pause. Let's refactor." But I do this because, again, years of building software, I know that it's just going to get a mess. I'm not going to be able to navigate. And for me, it's important that I need to understand my code and I need to have the structure for that. So I guess this doesn't change, and maybe the people who are vibe coding, they're going to come around to learning the same lessons that we all learned the hard way.
Yes, I guess. But how do you think AI is changing code security and also security in general? What impact do you see it's already having?
I mean, there's definitely a change, right? I think even for our security tools, there's a big change in the sense that it's very powerful and helpful. So even if you run deterministic algorithms like static analysis to detect issues, you can still enhance those deterministic algorithms with AI, right? So for example, we talked about the taint analysis. Your deterministic analyzer needs to have a lot of knowledge about all the libraries and frameworks that are there, and there are millions, right? And so AI can help you with gathering knowledge and information and then feed that into a deterministic algorithm, right? So you can combine technologies, and that's definitely changing, I think, the static analysis but also dynamic analysis and other security tool areas.
And then what we're also seeing is fixes work quite well. Like if you throw half a million lines of code into the context window, it's not working so well. But if you have a very specific task, if you say here's a deterministically found security issue, here are the 20 lines of code and that's the problem, then AI is very good in fixing those issues, right? So that's very helpful, because it's about fixing and not just detecting, and AI is super powerful here.
We also see a change in how code and applications are built, right? So if you think about applications, traditionally you have this backend, frontend, and in the backend is a database. If you remove that database, then you don't have a SQL injection anymore, right? But if you add an LLM to the backend, then you have a prompt injection, maybe, another vulnerability where the attacker can modify the system prompt or your prompt engineering and then mess with the LLM logic or the output. So the threat landscape changes and attackers adjust for it, and certainly the tools and the industry adjust to this, and that's maybe taking a bit of time, right, if you think about all the COBOL code we are still seeing.
Yeah, but I guess we can add prompt injection right up, we can pin it up there with SQL injection. In fact, who knows, prompt injection might become even more of a security issue.
Yeah, I think as text becomes code, right, I think prompt injection is kind of like the new code injection, right? The human language is the new code, and so if you inject human language, then suddenly that's your new code injection. So that's interesting from a security perspective.
And what about with coding assistants? Are you seeing things change in terms of how we think about code security?
I mean, I think the big problem in terms of security is that you produce code much faster, and writing code is not the challenge anymore. And so suddenly the new bottleneck is how are you verifying all that code, right? That's the new bottleneck: not to get your code done but to verify that it's actually secure. And if you don't, then that leads to security issues or quality issues, which then in the long run lead to security problems, right? So I think that's the big new challenge for security or code security: how do you verify all of that faster-produced code at scale and at speed?
And what is your take on what is working so far? I mean, the obvious thing that I'm hearing a lot of engineers and engineering leaders say is, well, we need to scale code reviews. We need to figure out ways where humans can look at code reviews, more of them, meaning let's add tools to them, let's add additional context. But outside of that, do you see some other maybe promising areas where we could actually verify, strictly from a security perspective, is this code secure?
Yeah, I think you mentioned already the key things, right: to add tooling to automatically verify code as you produce it. I think there are also areas where, and SonarSource is pioneering something here in the field, you basically look at how LLMs are trained and you make sure the data set that an LLM is trained on is actually free of common security issues, right? And if you do that and train your LLM on high-quality code, on high-quality data free of security or quality problems, you are producing from the beginning a much more secure code, and that's maybe another thing where in the future we will see more of this.
And speaking of this, again, because you see a lot of code, you do a lot of security analysis, do you see AI-generated code introduce different types of security issues than humans would, especially because we know that LLMs are trained on human code in the end?
I think they're doing the same mistakes, for the reason you mentioned. Maybe the prevalence of certain issue types changes, right? The issue types don't change so much, but what we are seeing, for example, slopsquatting is a good example, where AI proposes to use a library that doesn't even exist, right? And then an attacker can register in npm or Maven Central that non-existing package, and then with that you suddenly include a malicious package and there's the backdoor. And so this security issue was known before, and we had dependency confusion, but it's just less likely that a developer mistypes a dependency, while with AI that prevalence suddenly changes, right? There's an acceleration of that, while other issues maybe decrease, right? I could imagine, I'm not seeing this for now, but I could imagine hard-coded passwords, some issues that are just one-liner issues, maybe decrease a little because AI is able to learn that I shouldn't do that, and then we could see a reduction of those issues. Still, human developers can add them, but maybe the more AI-generated code is used, we see less of them, and then maybe we will see more of the complicated issues, right? Issues where you need to combine multiple code snippets with each other to form a security issue, that is then not so easy for AI to grasp. And so definitely some changes in the prevalences of what we know of already today.
What are some commonly misunderstood things about the security industry? Things that we could call fallacies.
I mean, I come from the security industry, right, and I moved more to the developer side over the years, I would say. And now, stepping a bit back for a moment and looking at the security industry, it's quite fascinating how we have this separated industry and community from the developer community, right? Where we both talk about code and bugs basically, but one side is maybe more about building things and the other about attacking and destroying, and so they are a bit distinct somehow and separated. That's interesting to me. And I think one fallacy that comes out of this is that the security industry is all fascinated about the security problems and then is selling products basically that promise you can have security just as a product, right? And I mean, there's a lot of money in that industry and it's driven by compliance and fear of data breaches, and so I think as a CISO you have a hard time knowing what product should I use and buy. And often I think a mistake here is to look at security as a product and not as something that you are building into the process of development, right? Because I think in reality that's what you must do, and not have a tool that finds you yet another 1,000 issues if you hit the scan button, right? But something that embeds into the process, finding issues but then engaging developers and helping you fix things. And I think that's the biggest fallacy to me.
We talked about the ownership that comes with that, maybe, right? I think there is a bit of this mysterium about security, that it can only be owned by experts who are top-notch hackers, etc., where I think the lines are a bit more blurry here, and it's more about fixing things and not just so much about the exploitation stage all the time that the security industry talks about and finds fascinating, and I myself am guilty here and find fascinating. Lastly, maybe this: there is no perfect security. If you get the understanding of the security industry, then maybe that's a fallacy here, that there's no perfect security, unfortunately.
Yeah. This thing about vendors selling products promising your organization will be secure, your code will be secure, and the fact that the reality is a lot more... it doesn't really work like that. Like, you need teams, you need people who care about it. Sometimes, I guess, you can have a team that uses zero security tools producing really secure code because they're just experienced engineers or working in the domain that they understand, and you can have the other way around as well: you can have a team that has all these scanners and whatever and their code is still not great, not good, and unsecure. It reminds me of the developer productivity term, like how productive are my engineers, and again there's vendors selling all these tools saying, hey, measure this and you will get this, and we just see the same thing. So I wonder if it's just a thing of there are just some things that are just hard because there's a lot of moving parts. You cannot just measure one thing, because we can optimize for that and still the outcome will not be great. I wonder if there's just some areas, developer productivity is one, security maybe; maybe software is just hard.
It is, right. I think the more complex software you write, and every developer knows this, right, the more hard problems you have, bugs you have, and also the more security problems you will run into, and that's just natural. We have a great vulnerability research team at Sonar, and those guys are picking the most popular open-source projects in the world that are deployed everywhere, with great communities, great maintainers, bug bounty programs where people get paid if they find something, etc. And it's fascinating to me every time they choose such a high-profile target that they go in and they still find something, right? If you're motivated enough and look hard enough, I think you can find something. And unfortunately that's the reality, right?
As a software professional, as a security professional in the field for 20 years, what can you advise to me as an engineer: how can I know that my software is secure enough, or at what point should I stop, and how would you think about this? Obviously there will be differences between if I'm a one-person tiny business, a midsize company, or a very large company. How would you advise engineers to think about good enough security? Okay, I can move on. This is good. Let's do the other stuff.
Yeah, it's tough because we said perfect security is hard. But then to the question, what is good enough and how can you solve this? I think using tools is the first thing you should use, right? I think it's a bit like securing your house, right? Like, you should make sure you shut your windows and doors and have some basic hygiene. It doesn't mean that a highly skilled or funded attacker can break in, but you can make sure you shut those windows and doors. I think with software the challenge is a bit you're adding new windows and doors every day basically, like with the new features you're adding, and so I think you need some automation for that to have your
basics, right. And then I would recommend basically you can start with an initial assessment: where am I standing today, right? Like, you can hire professionals or use a tool for this and kind of assess where do I stand today, what are my most critical issues I should fix, and get them fixed.
And then, more importantly, as you're adding features and as you're coding, making sure you're not adding more on top of that, right? Making sure you're not adding more security vulnerabilities and also you're not adding more technical debt and quality problems that in the long run lead to security issues. And I think here automation is key, basically.
And then after a quarter or something you can run that assessment again and look at where am I standing, and hopefully you have been very productive as a developer and adding new features that didn't slow you down, but also you increased your security posture at that point.
It's a never-ending story and it's a growing field, right? Like, we always need to be aware of the latest changes. Right now, LLMs and prompt injections: you'll probably need to ask yourself, if I'm building on top of LLMs or I'm invoking APIs, can they go in there? And then the next thing will come again, and the next and the next and the next. I guess keeping an eye on the OWASP Top 10 is never a bad thing, just to cover the very basics.
Yeah, I agree.
Now, as a closing question, I'm going to put you on the spot here. Which programming language do you think is the most secure? The one that you are very happy either using or observing, like, okay, this language itself seems to help prevent a bunch of security issues to start with.
I think the newer languages are more secure. I like Go; it's a good example. I think by default things are just... new languages learned from the past, from older languages, what goes wrong. But I think also other languages are evolving. I think Java, we see that a lot in enterprises, and I think it's quite secure to use. So that would be my answer here.
No, I like that you dropped Go. It's getting pretty good traction with startups as well, including now even for building web stuff. It's picking up. So I guess it's all down to people's tastes, but it's good to hear. So Johannes, this was really interesting. Thanks for coming on the podcast.
Yeah, thank you. My pleasure to be here and thanks for the invite.
Well, thanks very much for this.
Thanks a lot to Johannes for taking us deeper into the topic of code security. The thing that I found the most interesting is just how hard it is to define exactly what makes code secure, because there are simply so many impossible attack vectors, from using a dependency that gets compromised, to AI generating code with glaring security vulnerabilities like not validating inputs, to accidentally leaking credentials. The list just goes on.
Security feels like this invisible thing across software. As long as there's no security issues discovered, it doesn't get much attention. But once there is, then there's a scramble on what to do. As a professional software engineer, we need to keep ourselves up to date with common security vulnerabilities and how we can defend against them, including the new ones that AI tools introduce.
For more details on security engineering, see the Pragmatic Engineer deep dives linked in the show notes below. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show.
Article published
