Wie Anthropic Software entwickelt: Thariq Shihipar über Claude Code und die Zukunft des Engineerings

VideoCURTVortrag

In einem Interview mit Ryan Peterman erklärt Thariq Shihipar, Softwareentwickler im Claude-Code-Team bei Anthropic, wie sich Softwareentwicklung durch moderne KI-Modelle verändert. Bei Anthropic nutzen Entwickler Claude als Denkpartner auf höheren Abstraktionsebenen, während das manuelle Schreiben von Code im IDE in den Hintergrund tritt. Shihipar erläutert die wachsende Bedeutung von Harness-Engineering, autonome Workflows (Loop Engineering), Sicherheits- und Teststrategien gegen Systemausfälle sowie die Gründe, warum technisches Grundwissen trotz KI-Automation unverzichtbar bleibt.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Verändertes Onboarding: Neue Ingenieure bei Anthropic klären technische Setup- und Codebase-Fragen direkt mit Claude; menschliche Onboarding-Buddies dienen primär der sozialen und kulturellen Integration.
  2. Verschiebung zu Harness-Engineering: Bessere Modelle erfordern komplexere Ausführungsumgebungen (Harnesses) mit Sandboxing, Auto-Modi und Artefakten, um mehrstündige autonome Ausführungen sicher zu steuern.
  3. Wissensarbeit wird zu Code: Aufgaben wie Buchhaltung oder Videoschnitt lassen sich mit Claude Code, Python-Skripten und FFMPEG automatisieren, statt spezialisierte Desktop-Tools manuell zu bedienen.
  4. Herausforderungen bei Computer Use: Browser- und Desktop-Automatisierung bildet eine Zustandsmaschine, deren State Nutzer nicht voll kontrollieren; APIs und das Model Context Protocol (MCP) sind oft stabiler und schneller als visuelle UI-Klicks.
  5. Loop Engineering statt Mikromanagement: Systeme, die Claude selbstständig in Schleifen ansteuern (z. B. für Issue-Triage und Feedback-Analyse), erfordern robuste Verifikationsmechanismen und genaue Zieldefinitionen.
  6. Prävention von Produktionsfehlern: Da KI-unterstützte Teams Code deutlich schneller veröffentlichen, setzen Entwickler verstärkt auf Request-Replays, gemockte Datenbanken und Chaos-Monkey-Tests.

Warum das relevant ist

Der Werkzeugeinsatz bei Anthropic gewährt Einblicke in Praktiken, die sich bald branchenweit etablieren könnten. Die Verschiebung weg von manueller Syntax hin zur Orchestrierung von Systemen, formalen Spezifikationen und robuster automatisierter Verifikation zeigt, wie sich das Anforderungsprofil für Entwicklerteams grundlegend wandelt.

Einordnung

Shihipar stellt klar, dass KI-Modelle zwar Implementierungsarbeit abnehmen, dies jedoch neue architektonische Herausforderungen schafft. Insbesondere das sogenannte Harness-Engineering – also die Umgebung, Sicherheitsfilter und Schnittstellen rund um das Modell – wird zu einer anspruchsvollen Softwaredisziplin. Wenn KI-Systeme über Stunden autonom agieren, verlagert sich die menschliche Verantwortung auf exakte Zieldefinitionen, Fehlerisolierung und systematisches Testen. Wer Entwicklungszyklen beschleunigt, ohne gleichzeitig automatisierte Prüfpfade (Verification Loops) aufzubauen, riskiert eine Zunahme kritischer Systemausfälle.

Transkript

Vollständiges Transkript anzeigen (7.591 Wörter)
Thariq Shihipar: Most software we wrote before LLM’s was not very good. Ryan Peterman: This is Thariq. He’s an engineer at Anthropic on the Claude Code team and I asked him all about how their engineering team makes the most out of the models today. Can you move up like a higher abstraction level, you know, like, can you build the system that builds the system? Thariq Shihipar: Is there something that you feel is kind of the next big shift that is gonna diffuse into the industry? Yeah, I think there are a few different ones… Here’s the full episode. Ryan Peterman: My goal with this conversation is to ask you as much as I can about how can you get the most out of the models specifically for software engineering so that people in the industry can kind of learn from the best practices where Anthropic’s having success. And so to start the conversation off, I’d like to ask you, let’s say I was coming from a company that’s maybe less AI-pilled and I was onboarding onto your team, what are the most impactful things that you’d tell me to start doing to successfully onboard? Thariq Shihipar: I think the number one tip we have for you know, both people inside and outside Anthropic is that like if you treat Claude like a thought partner and give it like the context that you need, then you can usually figure out the next steps, you know? And and so I think that like starting with okay can Claude do it? If not, why not, you know? And uh, using that as the starting place, I I think is really important. And I think just like always thinking about like, okay, what’s you know, reflecting on, can you move up like a higher abstraction level? You know, like, can you uh, you know, set up a system? Can you build the system that builds the system, right, instead of just like building, you know, the the product itself. Ryan Peterman: I remember when we used to onboard people, you would get assigned an onboarding buddy, so someone who kind of really knows the codebase and you can ask all of your trivial setup questions to. It sounds like Claude fills in a lot of those gaps. And in practice, is it 100% you don’t need an onboarding buddy and you can just ask the model for everything? Thariq Shihipar: You do need an onboarding buddy, but more from like a, kind of like social and cultural perspective than a technical perspective, you know what I mean? I I think that like from a technical perspective, you can basically just like work with Claude and, you know, if you’re like good at it, like you can get what you need to onboard. But I think just like having the context of, you know, how to work with the team and also just like, how to like, you know, get buy in on what you’re building or, you know, understand like how the team works together, or just even like have a friend to work with, you know, I think is really important. So, yeah, we still have onboarding buddies, but not um, it’s not like nearly as much technical lift as as it used to be. Ryan Peterman: There’s a big difference between the external perception of the AI capabilities and the internal Anthropic usage because I’ll talk to friends and they say, yeah, what Boris is saying is the reality. And so, my question is, why is there such a big difference between that that internal perception and that external perception? Thariq Shihipar: I think we have a good track record, kind of of when we make these like claims, now you’re like, oh, you’re you’re, like it turns out to be true on to a long large enough time scale, you know? I think a lot of engineers are not really in the IDE anymore, you know, like, yeah, the code being written is like, I I talked to large enterprise customers where, like, you know, they haven’t like typed a line of code themselves in six months, right, so. We think of it part of our jobs, you know, to figure out how to work at a higher abstraction level. And like, if I spend like a day trying to figure out how to get Claude to like, you know, do this autonomously and I fail, that’s like fine, you know, and maybe even good, you know, cause now I can be like, oh hey, like Claude is not good at this, how do we make it better, you know? But I think if you’re working like an average job, you know, your job is to produce the output and like, I think automation always has a cost, because you’re like, it’s an investment and if it works then great, you’ll make, you know, a return on your investment over time. Um, but if it doesn’t, now you’ve like wasted all this time. And so, I think the nice thing about the models is that, you know, generally the chance of the investment working out is higher and higher because like the models are getting smarter and smarter. Uh, but you still have to sort of take that leap as an individual. Your boss is not going to be super happy with you if you’re like, you know, uh, if you’ve been working on your harness setup the entire time and not shipping code, um, but yeah, I think it’s mostly just like culture and thinking of it as our job to sort of live in the future and then we try and like build it into the product and our harness as well, so that you don’t have to do as much manual step, yeah. Ryan Peterman: Are there examples that you come to mind when you think of things you don’t see people doing on Twitter that are really having a big difference at Anthropic that you’d kind of recommend? Thariq Shihipar: Like, using Claude Code for knowledge work, I think is really valuable. The way a technical person does knowledge work I think is very different than the way a non-technical person does knowledge work, um, these days, because you can get the models to do it. And so, I think like, you know, increasingly, if you can figure out, like, hey, this is a task, the task is made up of code-like things, you know, and like, how do I tell the model the code steps to take, you know, in order to do it, right, uh, you can like, you can do a lot more, right? So I, I think I do like a lot of accounting with Claude Code using on the personal side, using like, uh, scripts and Python instead of Excel, you know? Or, like, I do video editing using like FFMPEG and these libraries to render visuals and things like that. And so, I think most of knowledge work is like, uh, reducible to code and coding agents, if you like think about it well, you know, and um, I think that that’s like one I think like big difference, uh, that I see like other people doing less of, I think. Ryan Peterman: We talked a little bit about the model and the harness, and uh, I want to ask you, what’s the relationship between the model and the harness in terms of getting the best end results? Like, obviously, the model’s important, so how important is the harness, and are there any examples you can kind of share of, Yeah, I mean, the harness is super important. I think you I I think that sometimes there’s this idea that the harness doesn’t matter because the models will get better and better, and if the models just does everything perfectly, then why do you need a harness at all, right? And I think that like, in practice what we see is like the models get better and better, and so the harness needs to become more and more complicated uh, to allow the model to do more things, right? And so, um, an example of this is, uh, auto mode. And so, auto mode, you know, is a classifier that runs after every, you know, uh, task that you normally have to ask a permission uh, prompt for for for Claude, and back when we were like Opus 4 or even Opus 4.5 maybe, like, it was not so bad to hit enter on the permission prompts because like the turns would only last a few minutes anyways, right? And now, Claude is running, you know, can run for hours, and so you really need that ability for it to, you know, do work safely, stick to your instructions, and auto mode is really complicated software, you know? Sandboxing is really complicated software. Um, and then like, you know, you also think about things like, okay, like Claude can do so much more work now, um, how does it represent the work that’s done? Like it’s worked for like eight hours and you want to know what’s done in eight hours, right? So this is what we use artifacts for, right? And artifacts themselves are a form of prompting, because like, how does Claude represent that in a useful way? Like there’s a lot of different ways that it could uh, represent that information. Um, and so I think roughly, what we see is like the harness, harness engineering is definitely like this like mix of science and art. I think it is very unintuitive in a lot of different ways, um, but I think it has like big abilities to like unlock new parts of uh, you know, like model behavior and uh, yeah, I just see the harnesses get more and more complicated and it’s kind of hard and harder to like actually vibe-cord your own, which is like a little bit unintuitive to me because of like how good the models have gotten, uh, but yeah, things like auto mode and workflows and things like that that which are like actually quite complicated pieces of software, become like really load bearing, I think. Ryan Peterman: As LLMs have advanced, more and more of the human part can be taken out of the loop. You know, at Anthropic, what percent of your or your team’s changes are fully autonomously made versus actually pairing with a model like people were doing more of like around a year ago? Thariq Shihipar: Um, I think it’s very dependent on, you know, what you consider autonomously made or what, what team or what function you’re working on as well. You know, for example, our designer might give you a figma file and then you pass the figma file to Claude Code. Uh, Claude Code is great at using the figma MCP, but like the designer has put a lot of work in there, um, I think that like roughly, um, our goal is to sort of get Claude being able to do all of the like, the glue work that ties everything together, right? So okay, you’ve got like a figma file, do I really need to like rewrite that in React or something like that? Probably not, right? Like I think like that work has been done once, right? And so the way I think about it is, what’s the unique work that I need to do everyday, right? And, uh, the more I’m like doing that unique work, the better. And and like I think there’s just a ton of demand for like, you know, unique work and thinking. And the more I’m doing something I’ve done before and like, okay like, can, can Claude do this? Or like if I’m just translating what someone else has done, you know, and like, oh like, can Claude do this as well. It’s really blurring the lines of like what, what does it mean, even, even if Claude has fully generated this PR, you’ve probably done a lot, made a lot of decisions and, and added a lot of context along the way, so. Ryan Peterman: How close are we to a world where someone just says, hey Claude, here’s the ticket, just don’t tell me until you’re done? Thariq Shihipar: Yeah, so I mean it depends on how good the ticket is. If someone has like perfectly written out a software spec for this ticket, then yeah, actually Claude can probably do that, right? But I think that like, um, to start, you know, like let’s say that we get a GitHub issue, that, you know, is this worth fixing? Is this a feature like, oftentimes issues are like also feature request or something or there might be like multiple things happening that maybe you combine together in a different way, you know what I mean? And so I think that there is a lot of that work is like, okay, what’s the vision for the product, where where do we want to go? How do we like, um, make sure what we what we’re building is cohesive? And so I think like for a specific spec, like a specific enough spec, Claude can do it. In practice, people don’t have haven’t figured out what they really want, you know? And haven’t figured out the unknowns or the shape of the problem and things like that. And maybe the way you would have done it before is like, you start writing the code and you’re like, oh like, what am I supposed to do now? Like or like, you know like you, you figured out that way. Uh, now I think you need new ways of figuring out what you don’t know yet, right? And I think you can still chat to Claude really, um, but I think, uh, that’s the hard part and I, I think definitely a failure mode is like, you you tag Claude, tag and you’re like, hey, please do this, you know, one sentence or less description, no previous context or memory, um, and then it does it and you’re like, oh no, I don’t like it, you know, and then you’re like, you know, just iterating forever on that, right? versus like figuring out what you actually want like really quickly. Ryan Peterman: When I was working at Meta, like the ambiguity of the task really ranges, oftentimes proportional to people’s, I mean levels aren’t everything in software engineering, but you know, it’s roughly senior engineers kind of take that ambiguous, uh, business need and convert it into something that’s a lot more concrete. And then people who are maybe newer in their careers, they kind of do the implementation work, um, and like intern projects were kind of almost line by line specified. If I was still uh, working at Meta and I had interns like just kind of give the, give Claude my intern’s back and it would kind of do it. What kind of work does an intern do then at Anthropic, if Claude can kind of handle it? Thariq Shihipar: What we see in practice is that there are just so many new types of work that no one has ever done before, right? And I think that like we all have to figure that out. And so, for example, how do you do an eval, you know what I mean, against like these new coding behaviors, right? Like how do you make sure like uh, that like, you know, how do you measure Claude’s performance across like millions and millions of users, across doing all sorts of task? Sometimes these tasks might have trade-offs, you know, so I think there are a lot of new work to be done, especially like in the research side of things and um, I think it’s harder and harder to write those specs, but there’s also less and less experience that is relevant, you know what I mean? And so like I think that, yeah, having some experience means that you can sort of, you know, you you know how to get work done and you know how to learn, but I think you also get those skills out of college and and you know, now I think as an intern, trying to figure out like, okay, what are the things that people have not done before and doing that, I think is really exciting. And I think that there is, um, yeah, I think internships are less, write the React code, you know what I mean, and there’s so many new problems to solve. We need people to solve them, it’s helpful if you have like fresh set of eyes on it, you know? Um, but it is definitely be more proactive and like, opportunistic maybe, than uh, than before where if you it was maybe more of like a pipeline, yeah. Ryan Peterman: We talked a little bit about knowledge work and that makes me think of computer use and browser use, and I want to ask you, you know, how far away is that from being widely adopted in the industry and impactful and maybe you can talk about how Anthropic uses this since it feels like are always far ahead. Thariq Shihipar: I think computer and browser use, the models have gone a lot better, Opus 5, I think is like a really good computer use model, but there are like these weird edge cases where like, for example, it can’t type a password on my behalf because like my passwords are in one password or something and you know, it can’t access that and so like it gets stuck and um, I think there’s still like those edge cases which are UXy edge cases. Um, and then I think also obviously computer use is also, uh, you know, the more and more people turn APIs and MCPs into like new ways to use Claude. I think that can also take a lot of the use cases that you’re using computer use for, right? And so going back to like, oh yeah, knowledge work, everything is code. I think like there are some things where there’s no API, there’s no way to execute in code other than just like opening up your browser, but I think increasingly there are more and more ways and I think CloudTag is a great example of like, we’ve just sort of like tried to roll up all these things into APIs that it can use, and so, um, it feels like it can do a lot of work on your behalf even the, even though it’s not literally running like a virtual computer and clicking things, you know? Ryan Peterman: I have noticed that whenever I use, uh, any kind of computer use tooling, it feels painfully slow. I look at the cursor and it’s sitting there for 10 seconds, moves over, you know, goes there for 10 seconds. What where’s all that latency coming from? Do you have a sense? Thariq Shihipar: I think that like it’s hard to make a small model that’s really really good at it, I think just because, there’s a lot of knowledge that you need to have about this task and how these things work together and stuff and so, uh, you want like a smart model and smart models, you know, take a little bit longer time, it’s like, I think a lot of people thought we’d get here sooner, but I think it’s just been a harder task than we expected, um, I think like one thing I’ve someone’s told me about computer is before is that like, it’s a state machine where you don’t control the entire state. If you are on the Doordash website or something and you want to add something to the cart, and you add it incorrectly, now you have this new flow to like undo it, you know, now you need to go click and now you need to go delete it and you can make a mistake along that side as well, right? Whereas like in code, you can sort of like undo, git, you know, whatever you control all of the state, um, but for computer use, like each action is if not irreversible, it’s like uh, you know, much harder to reverse than than than others. Ryan Peterman: Anthropic, um, I get the sense that employees have a lot of budget in terms of the compute to kind of speed up whatever it is they need to do. And so to if we were to spur the imagination of people who use the models to make them more productive, assuming they had infinite compute, like what what type of workflows would you start telling someone to do if they had infinite compute? Thariq Shihipar: I think this difference is slightly more exaggerated than you’d think, I think, you know what I mean? I think that like, for example, I use my max sub on the weekends, and I have almost never hit a five hour limit. I think the models are really smart and I think that like a lot of times when we’re spending a lot of compute, we’re just trying to find sort of capabilities or we’re trying a bunch of different things, and and it’s more about like us figuring out model possibilities, you know what I mean, than getting a lot of work done. When we’re testing for math, for example, we’re trying to understand how smart is the problem the model and that is useful to us, that’s useful output to work and it can sometimes solve like the Riemann hypothesis or something, or like not make progress, you know what I mean? People can replicate what we do at home just by thinking at a higher level of abstraction. For example, I’m trying to like get Claude to draft feedback for you, and so I want the funnel for like Claude has drafted some feedback to the user, has submitted feedback to be really good, right? And so I monitor that funnel, and then I ask Claude, I had some ideas, but then I was also like, oh what if I ask Claude to trying and improve the funnel? You know? And be like here are some ideas, can you come up with some as well? Let’s figure it out. All of those things you can kind of do yourself right now, you know, if you’re like, okay, let me when I’m making a feature, let me like, uh, annotate it with events, you know, let me make sure that Claude has access to those events, uh, let me run like a loop in the morning every day to check, you know, what events have fired and what changes were made, maybe even like let me proactively suggest some ideas, right? This can all happen, I think within a fairly reasonable amount of compute. I actually what I see more often is people running into limits where they’ve actually done kind of the opposite. They’ve started with a small mid scope task, you know, it’s like, oh, hey like, uh, you know, refactor this function in this way, and then Claude does it and, and maybe it’s like has some follow-on effects because like refactoring this has means you have to do some other work too, and that’s not exactly correct, or like you know, you’re iterating there and it’s like, you know, you just spent a lot of time, whereas uh, if you had sort of stepped up a level, told Claude your goals, then figured out like okay what are the details, you know, do some exploration, uh, maybe write out the schema or like, you know, and and then work with it, then let it run, you can probably get the same output. Ryan Peterman: I remember it was going kind of viral this idea of creating loops, um, and I mean that that does feel like a pretty, uh, expensive sort of thing to set up if you’re just asking it to kind of keep hammering away. So maybe first could you define this loop engineering thing and then I’m curious how often do you use it and you know, would you hit limits if you were on a max plan doing that kind of stuff? Thariq Shihipar: Yeah, so I okay, so loop engineering, roughly, it’s like uh, instead of prompting Claude directly, you’re setting up a system that prompts Claude, you know? And uh, you don’t have to think of it, if you use Claude Tag, a lot of this comes naturally, you just ask it like, hey every day do this thing, you know? And uh, that will uh, that’s a loop. You can definitely do that sort of work right now, but I think if you’re trying to set up like let's say like 10 loops or something, you know, that are like uh, triaging your feedback and implementing and things like that, we do that sort of work but we also spend a lot of time making sure our skills and things like that are useful, you know, and like are good at like triaging, uh, that we can have we’re good at verification and so that we can like make sure that the changes land, you know? And at that level, you know, what it means is like now we have someone monitoring issues that we just could never have kept on top of before, right? And it just like increases our software development velocity and it’s worth a lot of value to us. And so, yeah, I think that like if you’re kind of at the scale of like, okay, you want it to essentially be autonomously running, you know, doing a software engineering job, for a certain cases I think it can do that, if you set up the verification well, if you set up your skills well, if you give it the right data sources, but, um, it is a lot of work, right? And I think that like you have to make sure that that work is valuable to you in the same way that if you hire someone, you know, you have to make sure that work is valuable. Ryan Peterman: I’ve talked to some friends who work at big tech companies, like Google, Facebook, those types of places, and then, they as they’ve become more and more AI-pilled, one thing that people have noticed is there’s a lot more incidents or SEVs in in their usage. And it’s natural in those organizations because they they read the code less and there’s more code flying out, you know, what countermeasures have worked really well for Anthropic to prevent breakages given that the code velocity is so much higher? Thariq Shihipar: Yeah, I think this is something that like is a byproduct of moving faster sometimes and we have to figure it out. Like I think that like, uh, you know, I don’t think our uptime is exactly where we want it to be either, but also where as a company where a little like almost six years old, I think around, right? So it’s like no company has grown this fast before, and a lot of that is because we’ve been able to like create more products faster than ever before, right? And so you can use Claude to make your uptime better, right? And I think the way we think about this is like, really good, what’s the dream testing environment? The dream like, you know deployment environment, like, can you like, you know, take request and replay them across like, you know, mocked databases and fixtures across everything? Can you chaos monkey everything? You know? And so, yeah, I think there’s a lot of different ways of testing and verifying your code, and that I think is really valuable for maintainability, right? It’s like, um, just having like all these ways of verifying it, uh, and then, uh, yeah, of course like there’s just like the human element of like, where do you want your code base to go? If you know that’s the case, probably you start thinking about like, you know, how you’d replay and undo and redo becomes really important, right? Just like, that sort of stuff. And so, you the but if it’s single-player only, that actually matters a little bit less, you know, like, you, like the average user is not going to undo a hundred times or something, right? And so, you know, many like single-player applications, undo stop working, pretty quickly, like, you know, after like four or five times, right? Um, just because it’s not that important, but a multiplayer, the ability to compose different things together is really important. Or compose different operations together is really important. And so, um, that’s just an example, where like, if you know the direction of your codebase, there are things you care about and you want the models to know. And, and it’s like, I think, it's hard to make a small model that’s really really good at it, I think just because, there’s a lot of knowledge that you need to have about this task and how these things work together and stuff, and so, uh, you want like a smart model and smart models, you know, take a little bit longer time, it’s like, I think a lot of people thought we’d get here sooner, but I think it’s just been a harder task than we expected, um, I think like one thing I’ve someone’s told me about computer is before is that like, it’s a state machine where you don’t control the entire state. If you are on the Doordash website or something and you want to add something to the cart, and you add it incorrectly, now you have this new flow to like undo it, you know, now you need to go click and now you need to go delete it and you can make a mistake along that side as well, right? Whereas like in code, you can sort of like undo, git, you know, whatever you control all of the state, um, but for computer use, like each action is if not irreversible, it’s like uh, you know, much harder to reverse than than than others. Ryan Peterman: Anthropic, um, I get the sense that employees have a lot of budget in terms of the compute to kind of speed up whatever it is they need to do. And so, to if we were to spur the imagination of people who use the models to make them more productive, assuming they had infinite compute, like what what type of workflows would you start telling someone to do if they had infinite compute? Thariq Shihipar: I think this difference is slightly more exaggerated than you’d think, I think, you know what I mean? I think that like, for example, I use my max sub on the weekends, and I have almost never hit a five hour limit. I think the models are really smart and I think that like a lot of times when we’re spending a lot of compute, we’re just trying to find sort of capabilities or we’re trying a bunch of different things, and and it’s more about like us figuring out model possibilities, you know what I mean, than getting a lot of work done. When we’re testing for math, for example, we’re trying to understand how smart is the problem the model, and that is useful to us, that’s useful output to work and it can sometimes solve like the Riemann hypothesis or something, or like not make progress, you know what I mean? People can replicate what we do at home just by thinking at a higher level of abstraction. For example, I’m trying to like get Claude to draft feedback for you, and so I want the funnel for like Claude is drafted some feedback to the user, has submitted feedback to be really good, right? And so I monitor that funnel, and then I ask Claude, I had some ideas, but then I was also like, oh what if I ask Claude to trying and improve the funnel? You know? And be like here are some ideas, can you come up with some as well? Let’s figure it out. All of those things you can kind of do yourself right now, you know, if you’re like, okay, let me when I’m making a feature, let me like, uh, annotate it with events, you know, let me make sure that Claude has access to those events, uh, let me run like a loop in the morning every day to check, you know, what events have fired and what changes were made, maybe even like let me proactively suggest some ideas, right? This can all happen, I think within a fairly reasonable amount of compute. I actually what I see more often is people running into limits where they’ve actually done kind of the opposite. They’ve started with a small mid scope task, you know, it’s like, oh, hey like, uh, you know, refactor this function in this way, and then Claude does it and, and maybe it’s like has some follow-on effects because like refactoring this has means you have to do some other work too, and that’s not exactly correct, or like you know, you’re iterating there and it’s like, you know, you just spent a lot of time, whereas uh, if you had sort of stepped up a level, told Claude your goals, then figured out like okay what are the details, you know, do some exploration, uh, maybe write out the schema or like, you know, and and then work with it, then let it run, you can probably get the same output. Ryan Peterman: Your role at Anthropic’s really interesting because of the external visibility that you have and, um, I think a lot of people, when they give career advice, visibility’s a, a good thing, but they don’t have this level of external visibility and so, do you recommend to software engineers, like they should be posting on Twitter and X and you know, if so, what, what advice would you give in that sense? Thariq Shihipar: Generally, the thing I say to people is that you should share your work externally as much as you can, especially, um, I think within certain companies like you might not be able to, but like maybe you have side projects or something like that. I think just, um, before I joined Anthropic, what I would did was like I spent a bunch of time working with different companies, building stuff and writing about it and talking about it, um, and like, this was really valuable because it like increased my surface area of luck, you know? And so, I think that like, um, it’s really like the bar is much lower than you think. Like basically whenever someone asked me for advice, I’m like, okay, I’ll sit them down, I’ll be like, I know statistically, I I tell this people advice to a lot of people and almost no one does it. And I know that Jared, Jared’s post about the Riemann hypothesis had like, he was just like, keep going, we’re just believe in yourself. I think in this case, it was mostly just Jared saying it’s okay to use compute to solve this problem and I’m giving you permission to do it, you know? And I think that like, that’s not exactly the same as like, you know, I believe in you, you know what I mean? It’s really just like letting the model use compute. So, I think that that is something I would, I like to tell the models right now is like, okay, I think this is a hard problem, use sub agents, you know, use workflows, like if you need it, right? So, I always tell them to like use its own judgment, but I’m giving you permission to do this stuff, right? And I think that like, uh, you have to sort of remember that the models by default, you know, do what maybe the average user wants which is like, they want it to respond and start doing work as fast as possible, roughly like, complete the task but not spend like a crazy amount of compute on it, you know? And so, I think that like, um, you have to sort of, if you want the model to do it differently, you have to nudge it slightly, right? So you might have to be like, okay, hey, like, I don’t want you to do any work yet. I want you to brainstorm, you know, I want you to like, think with me, right? And if you prompt it that way, it will start doing that. If you want it to spend a lot of compute, if you’re like, hey like, you know, sometimes I’ll say like, yeah, hey I think this is a hard problem, uh, feel free to use workflows, if I’m running overnight, I might just be like, hey, I’m going to sleep, you know, set a slash goal or something and then uh, let it let it run, so, um, yeah, I think there is a uh, just like giving it permission to do the thing you want. Ryan Peterman: When you recently removed so much of the system prompt, how did you prove that the end result was better? Thariq Shihipar: We have a bunch of user metrics just like how, you know, how much do people like the output of Claude, do something you get that survey, and you see it, um, we run evals against, you know, our internal eval, uh, and external evals to see like how it performs at these different tasks. Um, but I I think it is hard like sometimes you don’t realize that they’re not, if Claude is telling the user, if it does all its work and then is like, hey maybe you should go to sleep, there’s no eval for Claude tells you to go to sleep, you know what I mean? And now we’re like have to like catch this new behavior. Um, so it is hard, I think we spent a lot of time basically just, um, you know, like removing lines of the system prompt, running evals, uh, seeing how it works, seeing how people reported it internally, and then like adjusting, but it was like a full-time job for several people over long periods of time, and so, I don’t think I necessarily recommend everyone do this. I think that’s kind of why we wrote that post about what we learned from like adjusting the system prompt, and we think that that’s pretty general, so hopefully you don’t have to like now go through this like crazy iteration process. Ryan Peterman: I saw this tagline going out a lot with some of the stuff that Anthropic’s been, um, maybe like I think it was like Boris went on some podcast and the tagline was coding is largely solved. You know, should people still learn to code if coding is largely solved? Thariq Shihipar: I think being technical is really, really important. Knowing how do computers work, how does like, um, yeah, how does how do computer programs work, how do languages work, like what are the hard things and like what’s a back-end service like, what’s a cache? Like what’s what like, like you know, what like what is memory allocation? Like all these things are actually kind of really important to learn. Um, I do think it’s hard to motivate yourself sometimes to do, in the same way that like, you know, doing math by hand was not that motivating to me, you know, but some people just love math and did it. Um, I I think that like that’s probably something that people have to figure out, but I think it is really worth it. Like being technical is really, really important. Like we talked about at the start of or like earlier, where like the only way you can tell you solved, you know, the Riemann hypothesis or like made progress or whatever is like if you’re a great mathematician, right? And, in the same way, the only way you can tell if you’re like built great software is like if you’re a great software engineer, right? And, in the same way, how do you do that is hard, but like we’ve talked about some of this stuff before, just like staying in the loop, like putting like reflecting on your process and, and getting better, and, and you can use Claude to learn as well, and, and sort of explain things to you, so like treating it like a thought partner. Um, but learning like truly, I I think you know, Kaparthy says like learning should feel like effort, you know, and I think that’s like one of the hard things is like, a lot of times, even if you ask Claude to explain something to you, you might just like nod along and you’re like, oh yeah, like I learned it, but you didn’t really because you didn’t put any effort in, right? So I think that like, it is really technical. You should learn. It’s hard to learn. And then, sometimes like what school forces you to do is to learn that. Um, I don’t know truly like I’m not learning programming from scratch. So I don’t know exactly what how to do it now. I think if short of better ways, I would still like type out and build programs and run them and learn them, you know what I mean? Um, there might be better ways that are more cloud informed as well, but I think it’s really important. Um, I think when Boris says coding is solved, I think it just means like, um, you know, we don’t get stuck in the same ways that we used to before. Like I think coding used to be this very high variability thing where you’re like, oh like, could this bug take a day? Or could it take two weeks? You have no idea sometimes, you know? And I think like, um, on the whole, coding used to be, like, in real terms, like something that was very rare for something to go well. You know what I mean? Like, very few people in the entire world could write software and they were very, very rare. And even if you got them all together, there were so many other reasons why it wouldn’t work, right? And, coding was this like one of the rarest things in the world where like the chance of software going project going well was like, on absolute terms, very low, you know? Uh, and if you’re like hiring someone to like make software for you for something that’s not like a huge product, like if you’re hiring someone to make software for your car dealership, you were almost certainly not going to get the software you wanted, you’re going to get like essentially scammed, you know what I mean? Um, not because anyone was trying to scam you, it’s just like software is really, really hard and you know, it could only be spent on like the most important scalable things in the world, and now that coding is solved, I think what it means is that like you can use coding to do all these other things that we’ve not done before, but that’s not to say that like there’s not a lot of work still. Um, it’s just like it’s not this like incredibly rare, difficult thing that mostly like just doesn’t work and you have to spend like eight hours a day locked in to do. Ryan Peterman: Also, well, thanks so much for your time, Thariq. Really appreciate it. Thariq Shihipar: Yeah, of course. Thanks, Ryan. It was fun. Ryan Peterman: Hey, thank you for watching this podcast. If you liked it and you want to see the show grow, please support with a comment or a like. Also, if you have any recommendations for people you want me to bring on, please drop a comment. Guests like Barbara Liskov, Mike Stonebraker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I’m working on building the ergonomic keyboard that I wish existed. Here’s a glance at the prototype. It’s a split keyboard. So there’s two sides, um, this is in the case. But yeah, we launched on Kickstarter and we hit our goal within eight hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units. Um, we’re now working on the long journey of building the tooling now. And so if you still want to pick one up, I’ve left the late pledges open on Kickstarter, so you can grab one there. I’ll put a link in the description. Thank you again for watching the podcast, and I’ll see you in the next episode.

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Artikel:Ryan Peterman

    Wie Anthropic Software entwickelt und wie sich das Engineering verändern wird

    In dieser Podcast-Folge von 'The Peterman Pod' spricht Ryan Peterman mit Thariq Shihipar aus Anthropics Claude-Code-Team darüber, wie Anthropic KI-Modelle intern für die Softwareentwicklung einsetzt und welche Veränderungen der Softwarebranche bevorstehen.

    KI & AI· Diskussion

  • Artikel:Anthropic

    Power-User-Tipps für Claude Code aus dem Anthropic-Team

    Das Entwicklerteam von Claude Code bei Anthropic teilt bewährte Workflows zur Steigerung der Entwicklungsgeschwindigkeit. Im Mittelpunkt stehen parallele Ausführung über Git-Worktrees, strukturierte Planung, iterative Fehlervermeidung via CLAUDE.md und automatisierte Verifikation.

    KI & AI· Sammlung

  • Video

    Video:Edward Donner

    Der Agentic Development Life Cycle: Softwareentwicklung im Zeitalter von Coding-Agents

    Edward Donner stellt den 'Agentic Development Life Cycle' (ADLC) als Weiterentwicklung des klassischen SDLC vor. Anhand seines 40.000 Zeilen umfassenden Testbed-Repositorys 'Bench' gliedert er den Entwicklungsprozess mit KI-Coding-Agents in drei Kernbereiche: Harness, Handoffs und Humans.

    KI & AI· Anleitung

  • Artikel:Thariq Shihipar

    Neue Regeln für Context Engineering bei Claude 5

    Anthropic-Entwickler Thariq Shihipar beschreibt, wie das Team den System-Prompt von Claude Code für Claude 5-Modelle um über 80 Prozent reduzierte. Moderne LLMs wie Claude Opus 5 und Claude Fable 5 benötigen weniger starre Regeln und Beispiele. Stattdessen setzen Entwickler verstärkt auf sauberes Schnittstellendesign, progressive Offenlegung von Kontext und dynamische Referenzen.

    KI & AI· Meinung

  • X-Post:rari

    Harness Engineering: Wie robuste Umgebungen für Multi-Agenten-Systeme entstehen

    Ein Leitfaden und Video-Workshop von rari analysiert das Konzept des Harness Engineering für KI-Agenten. Wenn Agenten scheitern, liegt das meist nicht an den Gewichten des Modells oder schlechten Prompts, sondern an einer unzureichend definierten Ausführungsumgebung. Anhand des Claude Agent SDK und Beispielen von Anthropic und OpenAI wird gezeigt, wie Werkzeuge, State, Sensorik und Leitplanken strukturiert werden müssen.

    1960Lesezeichen142.239Aufrufe

    KI & AI· Sammlung

  • Link:Anthropic

    Anthropic Claude Code: Agentisches Coding im Terminal und in IDEs

    Anthropic stellt Claude Code vor, einen KI-Coding-Agenten für Terminal, IDEs, Web und Slack. Das Tool liest Codebasen per agentischer Suche ein, führt Multi-File-Edits durch, führt Tests aus und erstellt Pull Requests direkt aus der Arbeitsumgebung heraus.

    KI & AI· Tool

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.