Cognition-Präsident Russell Kaplan über Coding-Agenten, Benchmarks und Modell-Routing

VideoLangChainDiskussion

Im Gespräch mit Harrison Chase (LangChain) erläutert Russell Kaplan, Präsident von Cognition, die Entwicklung des KI-Softwareentwicklers Devin. Er beschreibt den Wandel von reinen Modell-Fähigkeiten hin zu Kosten- und Latenzoptimierung (Intelligenzsättigung), neue Bewertungsstandards wie FrontierCode und die Architektur hinter Devin Fusion.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Devin startete im März 2024 mit 13 % auf SWE-bench; im Juni 2024 wurde der Agent zum aktivsten internen Committer bei Cognition, bevor Großkunden ihn für Refactorings und Migrationen nutzten.
  2. Viele Programmieraufgaben erreichen laut Kaplan eine 'Intelligenzsättigung', bei der alle aktuellen Modelle ausreichen und Entwickler primär auf Kosten und Geschwindigkeit achten.
  3. Für Sicherheitsanalysen kombiniert Cognition GPT 5.5 (hoher Recall) mit Fable-Modellen (hohe Präzision) im selben Workflow.
  4. Da SWE-bench gesättigt ist, führte Cognition den Benchmark 'FrontierCode' ein, der in Zusammenarbeit mit Open-Source-Maintainern Kriterien wie Code-Mergeability, Scope-Disziplin und Stil bewertet.
  5. Das freie Auswählen von Modellen im Cloud-Agenten hält Kaplan für einen UX-Fehler; Devin Fusion setzt stattdessen auf intelligentes Routing und eine Sidekick-Architektur für 35 % bessere Preis-Leistung.
  6. Cognition trainiert spezialisierte Post-Training-Modelle (wie SUI 1.6 auf Opus-4.6-Niveau), die auf Cerebras-Hardware mit rund 950 Tokens pro Sekunde für Review- und Linting-Workflows laufen.

Warum das relevant ist

Unternehmen stehen vor rasant steigenden Token-Kosten, die in einigen Organisationen bereits an Personalkosten heranreichen. Wenn Standardmodelle für alltägliche Entwicklungsaufgaben genügen, verschiebt sich der Wettbewerb von roher Modellleistung hin zu intelligenter Orchestrierung, Sidekick-Agenten und spezialisierter Inferenz-Hardware.

Einordnung

Kaplan beschreibt eine spürbare Reifung der Agenten-Branche: Reine Benchmark-Rekorde weichen praktischen Kriterien wie 'Mergeability' und Wartbarkeit von generiertem Code. Die Einführung von Sidekick-Modellen und automatischem Routing zeigt, dass Agenten-Entwickler zunehmend die Komplexität der Modellauswahl vor dem Endnutzer verbergen müssen, um Fehlbedienungen und ausufernde Kosten zu verhindern.

Transkript

Vollständiges Transkript anzeigen (25.771 Wörter)
Harrison Chase: I started my own machine learning career at Tesla on the autopilot team. And I talked to my friends at Tesla today and the bottleneck is no longer just training bigger and bigger models. It's actually running the evals. Today I'm talking to Russell Kaplan, President at Cognition, the company behind Devin, an agent that went from viral demo to deploying code inside some of the most complex orgs in the world. Russell Kaplan: We essentially have an evaluator agent that can take a session and say, was it productive or not? It gave us the confidence to actually go to our customers and say, we are actually going to make a $10 million productivity guarantee. Harrison Chase: He explains why and how proactive agents are giving human engineers outsized leverage. Russell Kaplan: You have like all these great suggestions of fixes that need to be applied. Oh yeah, that looks good. Oh, I want you to change that here. Individual developers have essentially realized I can be the CTO of an army of 10,000 agents. Harrison Chase: We get into what running agents at scale actually costs and how Cognition drives it down. Russell Kaplan: There's organizations where the per person token spend is starting to eclipse the human salary spend. By being a little bit more clever about some of the routing, we can get about 35% better price performance with actually a slight increase in quality. Harrison Chase: And he argued there's no such thing as the best model anymore. Russell Kaplan: The Fable class models are really good for high precision, but we actually find that, GPT 5.5 and 5.5 cyber are better on recall. You have to use both. Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. Harrison Chase: You guys had a massive launch about two and a half years ago, three years ago at this point, and I think you pioneered a lot of really interesting concepts, especially around UX and interacting with Devin in Slack. I think that was the first time I saw that so prominently. So, I remember a lot from that launch, I remember a lot about the UX, I'm sure others do as well. What should people know about what you guys have been up to over the past two and a half years? Russell Kaplan: Yeah, we launched Devin in March of 2024. was the original demo video that went super viral. And at the time I would say, coding agents were just at the edge of possible. You kind of squint and see, okay, this is going to work at some point. And we sort of put together a first pass example of what could that look like, you know. Harrison Chase: Sweet, bench was like 13%, I think you guys hit, and it like tripled the previous best. Russell Kaplan: Totally. It was like, yeah, we were like, oh 13% sweet bench. Um, this is, you know, really exciting. And it actually that, that by the way, that, that number reminds me a lot of, some of our more recent evals and where we are now in agentic coding with a, with a new eval we put out, um, called frontier code, which we can talk about. But, uh, when we, you know, we kind of launched in March of '24, we had this point of view that eventually, you're just going to have, uh, agents as teammates that you can delegate complete units of work to. And what we learned is that took us a few months from, okay, this is a prototype that we can kind of see the future of to, this is something that's actually useful internally. I think it was June of '24 that Devin became the number one committer to Devin, which was like our first big milestone. Harrison Chase: That's still a while ago. Russell Kaplan: It was a while ago, but it was like a lot of manual dogfooding and like really kind of grinding to make the internal workflows nice. And, and then it took us a few more months to actually get this deployed in production and useful at a customer. And in, in the sort of 2024 era, async cloud coding agents, they really couldn't do most of the tasks in software engineering, but there were some, you know, niches already. And I think like one of the best early ones was migrations and refactors. If you ran essentially a, you think of like a regex plus plus style workflow across a large codebase with sort of bits of intelligence sprinkled in, that actually already worked reasonably well in late 2024. And so that's where we got our, you know, early product market fit. It was with bigger companies, you know, like more like enterprise companies who they just had lots of code that needed all this transformation. Harrison Chase: And why, why was that such a good fit? Russell Kaplan: There's a few like technical reasons that made sense. One is, um, for a very large refactor or migration or, you know, ETL transformation, whatever, it's sort of worth it to put in the effort to carefully prompt engineer your, uh, your cloud agents to be accurate. And so, you know, if you, if you kind of iterate on your prompt and your setup and your, uh, you know, the context you're feeding in and you, you get it like just right, and then you can apply it across, you know, 10,000 modules in your codebase, this is actually really high ROI. And you could, it's obviously much better than just sort of a find and replace, you know, string change. But um, it didn't have to have this like full general software intelligence that we're now delegating all sorts of coding tasks to. So that worked well, even in the early days, but it was, it was still kind of niche. And then in, yeah, December of '24, we launched Devin, so anyone could sign up. And we had this big debate internally of what should be, you know, what should we really emphasize or what should be the focus. And I remember we did like the, we just, we flew the whole company to Utah to just like lock in in December and get this out, uh, you know, kind of before the end of the year. And what we settled on was actually Slack as the primary interface. And so the entire launch video for Devin available self-service, uh, we called it internally at Devin. It was just @devin, @devin, @devin and really trying to drive home this sort of user experience change, which is collaborating with agents more like teammates. And so from there, we got a lot more users, we got a lot more traction. You know, cloud agents have just gotten better and better since then. Both as the infrastructure has gotten better, as we've matured on that, as the models have gotten better, as the harness has gotten better. And now I think most of the work, uh, is being done by just delegating to, to async agents. Harrison Chase: Yeah, maybe diving into that a little bit. What are some of the things that have gotten better that have allowed them to really take off? Russell Kaplan: So a few things. So I think people always talk about the models getting better and that's, that's super important. But I think the the kind of infrastructure maturity is What was what was the first model that you felt was like really actually good enough? Russell Kaplan: Oh, yeah, it's, you know, we kept saying, oh, this is the model, this is the model, you know, this is the new model. There were a few like phase changes I would say. I, I think sort of like the, the November of 2024, uh, like model series was, uh, one kind of that was like a big step function where it was like, okay, this actually work a lot better now we, we did a self-serve launch, you know, in part based on that, um, from OpenAI. And then, and then, you know, the Anthropic model started getting really good. Um, and we saw I think probably like, you know, July of '25, another big step change. And I think now we see, you know, with like Fable 5, this like new generation of models, continuous step changes. And now the step changes are so high, there's actually no longer about the capabilities increases. I think we're kind of seeing the opposite trend in some sense, which, uh, um, more and more of the tasks in software engineering are getting intelligence saturated. So like, literally all of the models can do them well. And, and so the, the sensitivity of a lot of developers has totally shifted from, wait, I need to be using the best possible model to holy cow, I'm spending so much money on my coding agents. How can I, how can I be a little bit more efficient, you know, about this? And I think, I think this is an underappreciated aspect of continuous gains in frontier intelligence, which is that, you know, for some workloads you always want the most frontier intelligence, um, especially if it's an adversarial game between parties and you know, your intelligence has to be higher than the the counterparty's intelligence, like if you're trading in competition or something. But for a lot of software you want to build, you know, there's kind of this, um, this saturation threshold where once it's good enough, what you set what you, what you care about is speed and cost. And I think that's actually where we see a huge amount of demand from the developers and the companies we work with is, okay, this already works. How can I now optimize the speed and cost? Harrison Chase: On the model side, maybe one question. What are tasks where you think Fable type models are like, where, where you see that jump today? Like where, where would you recommend people use them, where do you see people using them? Russell Kaplan: So I think we're seeing, um, maybe in like the, this most recent generation of models and Fable being one example, um, particularly big gains is in, first of all, there's like detection or mediation of cyber security vulnerabilities and we, there's a lot of guard rails now in the public versions of these models and so it's actually can be tricky to to get them to help. We see for example that if you're trying to go just like clean up my security backlog, which by the way is a huge workload right now for a lot of the a lot of the companies we work with. You know, the, the, the Fable class models are really good, um, for high precision, um, but we actually find that, um, you know, GPT 5.5 and 5.5 cyber are better on recall. Um, and so the best possible harness for finding and remediating security vulnerabilities right now, you have to use, you have to use both and you sort of filter down, you know, the the ones that are maybe better found by the GPT series with, with the Fable series. I think that's one big example. Um, another one is that processing data from integrations, for example, from like DataDog or, um, or other systems, it's, we do see like a step change in our evals in, in Fable in particular. Harrison Chase: Is that because that data is like so large and messy that it needs more intelligence to? Russell Kaplan: I think a lot of yeah, it's just like there's like a volume consideration, there's sort of a multi-step reasoning consideration and I think there's just you can tell that there's a lot of RL going on on like these very realistic data sources that have been scaled up a lot. Maybe the third category that's maybe the most noticeable is a category that we've only even gotten better at detecting and understanding recently ourselves with the work we've been doing on new evals for coding because to your point, when we launched Devin, and it was, you know, uh, in the teens on speed bench, now speed bench is totally saturated and so how do we find, how do we find the next, um, level of difficulty for evaluating models. We really, we we looked for all of the evals and we just we couldn't find any. They would actually kind of match our sort of hazy internal intuitions of this model feels a lot better. And when we sort of dug into it and really ask ourselves why is that, I think the core gap for evals that we found was around mergeability, which is, okay, this this code is technically correct, but like, would you actually merge it? Like, would you feel happy? Would this improve the quality of your codebase? And there's lots of different subtle examples, um, of what that could mean. It could be stylistically following the expectations that you already have, it could mean that it's done in a way that it's easy to modify in the future or that if other modifications happen, it gracefully handles those modifications. All these little stylistic things, which we felt like were not being captured in, uh, in the evals. And so, we we set out to basically make a new eval to, to measure this better and also just have a new higher watermark of difficulty for frontier coding capabilities. And then we released it recently, um, it's called Frontier Code. And it's once again, uh, it's it's it's the new hardest coding eval, uh, that the frontier code diamond subset is like in the teens, uh, you know, pre Fable 5 and now I think Fable 5 has, has pushed it to the 30s. So we have, I wonder how many more months we have before that eval is. Harrison Chase: I was going to say, yeah, and like a year, we'll be saying, oh, remember when it first got in the teens for this one? Russell Kaplan: Yeah. I'm wondering, I do wonder, like, it's going to get harder, you know, it gets harder and harder to make these evals. I mean, I started my own machine learning career at Tesla on the autopilot team and this was in like 2017. And the bottleneck was very clearly, you know, GPUs and compute, and it was like, okay, how do we get more compute, we just need more compute. And you know, I talked to my friends at at Tesla today and the bottleneck is actually no longer just training bigger and bigger models, it's actually running the evals. Because for a self-driving system, the interventions are so rare that you have to do enormous volume of driving to even find any problem at all in the stack. And, you know, I think for, for software engineering, we're not there yet, like we still find bugs, but, but we might, we might hit that threshold sooner than we think. Harrison Chase: We talk with a bunch of teams around creating evals for their own systems using LangSmith and other things like that. How did you guys create Frontier Code? Like what did that process look like? Russell Kaplan: So it started with, okay, we need to go find the sort of peak, high taste developers who are going to have really strong, both kind of stylistic quality, but also code correctness standards, to help us answer the question, would you actually merge this code? Not just is this code passing tests? Is it, is it functionally correct? So that's kind of part one. And so we actually did like a, you know, a deep collaboration and recruiting campaign with a lot of, um, the sort of leading open source authors who were, you know, really grateful for their collaboration on this. It wouldn't have been possible without them to go find, you know, what are the best known libraries in open source with, with really high coding standards. And then collaborating closely with them to helping code their, their human intuition of what makes a PR something I would accept, uh, into these very carefully designed tests. That was part one. Harrison Chase: What, what did those tests look like? Are they LLM as a judge? Are they programmatic assertions? Russell Kaplan: Yeah, so multiple. So, so part of them are programmatic assertions. Of course, there's like correctness tests, programmatic, there can be programmatic assertions on, um, kind of stylistic elements, too. Like, uh, okay, if you make this change in this pull request, uh, you know, you have to actually use this module, even though using this other module would be equivalent right now. Over time, if they drift, you know, if you didn't use the correct module, then like, you're going to introduce a bug in the future. So it can be asserted deterministically, but it requires the judgment of, you know, the human, um, author to settle. And then the other big thing, which I think is still underdone by people practicing machine learning and working with agents is, every single researcher on the Cognition team, you know, hand contributed and reviewed and eval the evals directly. And so, this is not, this is not something you can sort of throw over the wall and say, okay, the data quality piece, that's that's such a slog, um, someone else can do that. You, you sort of have to work on it directly too. Harrison Chase: I'm assuming they worked with the open source authors to put together these assertions and, and they would be running on these open source code bases, I guess? Like that was kind of the setup. Russell Kaplan: It would be on on the open source code base and then you know we could run it in our infrastructure. But then, um, we test and, and, you know, to avoid contamination, we haven't published the full set of questions. We published some example questions, but we're trying to preserve this eval for the community for as many months as possible until it gets, you know, completely saturated, um, but running, running, um, yeah, essentially running the code that already exists with the, the patches applied by these models. Harrison Chase: How do you guys score these evals? Is it binary? Like 01? Is there some numeric component, like if it, if you have 10 assertions on one of the tests and it passes nine out of 10, how is, how is that scored? Russell Kaplan: Yeah, so, uh, we broke it out into two components, uh, to have first, there's a binary element, which is essentially, yeah, did this pass all of the blocking constraints of how we would evaluate this PR. You know, the simple things of like, do the tests pass? Are there hard to are there hard deterministic criteria that the, um, maintainer of the code base has inputted must be met by the, uh, by the agent who's writing this code. Harrison Chase: Do you know off the, like, is that like five assertions or like 50 assertions? Like... Russell Kaplan: It's highly varied per task, but each task is, you know, like hundreds of hours of work of people putting into. So it's, it is like, this is not like a quick thing. It's the, the, the process of constructing a single test, uh, you know, eval line item in a dataset like this, it's, it feels like a, it's like a full project and labor of love per question, you know, where you're really trying to think holistically. What are all the things that I, as the expert maintainer, want to sort of imbue in my expectations? So that's how you get some of the binary pass/fail metrics. But then we also, to your point, that's often not enough to, to really know, okay, would I merge this code? And then also like, how would this code rank to someone else's code? Because there might be, you know, two different pieces of code you would merge, but one is like a little bit more preferable than the other. And so, uh, we also have the concept, concept of a score, which is essentially a linearly weighted aggregation of all of the non-blocking evaluation criteria. So if it's blocking evaluation criteria, we'd say, okay, you know, if it doesn't pass, you fail. But, like, uh, or, or this is just, you know, we're not we're not merging it. If it's non-blocking criteria, uh, for example, there's stylistic elements. A common one for us is, is scope. So I think one of the code smells of LLM still is like, they'll make the right change, but then they'll kind of mess with other files too that you really wish they didn't mess with. And so, that's not going to impact your correctness score, but it's certainly going to impact this sort of stylistic elements, um, and that's going to downweight, you know, if you had unnecessary touches in other files, that's going to downweight your aggregate metric. Same there for, you know, LLM as a judge, having heuristics that you impose in LLM as a judge can get added to this, um, to the, to the linear score. We also found that, um, reverse classical evaluation is is really helpful. So, commonly people say, okay, we need to make sure that after you accept this code change, the test pass, right? But we also care just as much that, um, you know, without this code change, the test should fail, right? And that, and that if you make this other code change, the test should fail. So how do you, how do you impose kind of blocking constraints on both sides? But yeah, evals are really hard. You know, we spent a lot of time on this. I think, I think we got it, we, uh, you know, we did it right enough that all of the current LLMs, uh, are pretty bad, are pretty bad at this eval and that also, more importantly to us, you know, as new models come out and we test them and we see their scores on Frontier Code, it's, it's roughly matching the vibes. Like, like if a model's scoring really well on Frontier Code, then we get really excited. Harrison Chase: I like this idea of kind of like binary pass/fail and then some numerical after that for the or or more explicitly I like the like blocking and non-blocking. So, so we're building a benchmark of our own, we call it like issue bench for Langsmith engine, which goes through and finds issues, and there's a bunch of stuff that like, I would classify under like the non-blocking stuff. Like it's, yeah, it would be nice if it did this, it should probably do this. And then there's some other stuff that's kind of like more blocking. Does it just find this like really bad issue should be more of a blocking thing. So I like that kind of like dual juxtaposition. Do you guys use Harbor as a format for running evals or do you have your own internal kind of like eval runner? Russell Kaplan: We do use Harbor. We we think in general, a lot of the standards are quite early and immature. And I think one of the funny things if you sort of look in the whole Devin codebase is, because it was the first coding agent, there's actually a whole bunch of stuff that like there might be a standard now or a correct way of doing things, and we just we rolled our own our our old implementation. You know, as a funny example, funny example is even a basic agent things, like the concept of, you know, skills.md file. Like, we we hit implemented a concept in Devin of of knowledge that, uh, was like before any open source standard existed for skills. And now we've basically grafted it on how can you also work with the open source standards, but it is like a recurring theme in our codebase that we've gone off and, sort of, invented something weird and then it becomes some some permutation of it becomes an open source standard then we incorporate it back in. Harrison Chase: We went down this rabbit hole talking about, uh, models and talking about the best models, but you also mentioned kind of like cost and and presumably kind of like speed and latency are becoming other issues. You guys have done a few things here if I'm correct, you have Devin Fusion, you also have your own series of models that you guys have post-trained. How do you guys think about this, uh, this section of the model universe? Russell Kaplan: We're kind of in this unique time in history where anyone can hire as many AI agents as they want, usually, in the way the tools are work. So, you know, uh, me as a developer, yeah, I'm not going to, I'm not going to use the cheap, the cheap model. I want to, I want to use the best one for everything. Um, but everyone is kind of collectively making that decision and then you sort of roll it all up and you realize, oh wow, like we are spending a lot of money. You know, there's organizations where the, the per person token spend is, is very rapidly approaching or even starting to eclipse the, you know, the the human salary spend. And so it would be, kind of crazy, you know, if, if like the way you ran LangChain for example, is anyone could just hire 1,000 people tomorrow without, you know, without talking to you, but that's kind of how we run our teams today. And this is starting to become like a real issue for us. And like I like to think that we're pretty like, you know, we're still startup, we're a lax like, AI-native kind of for organization, but we are absolutely caring about, like, token spend. I'm, like, are you, are you guys internally, kind of like worried about token spend for your own, kind of like engineering teams as well? Russell Kaplan: We have very high, uh, you know, compute investment in general. So I would think our internal spend, while being very high per person, to us, it's like, really valuable dogfooding and investment in everything we do. But our customers are definitely thinking about, if you just extrapolate the trend line, you know, it's going to eclipse the whole economy, you know, that that long with the, with exponential growth. And so, so, but so people are asking the question, okay, well how, how should we approach this? How should we even be thinking about this? And I think we now, there's two different trends that are happening at the same time that sometimes get conflated. So, so one is, obviously the models are getting a lot better and more and more capable. What's happening is, you know, if you look at the most frontier capabilities and then maybe the set of models that are just behind them or a little bit behind them, every time we have these new generations of models, we, we both move the high watermark on the frontier, but then this set of tasks that all the models can do is growing a lot. And the distribution of tasks that people are like trying to accomplish that are any in their everyday lives are changing a bit, but actually not that much. You know, and so it's like, okay, I'm trying to build a frontend, you know, for my marketing registration page. It's like, that's not that hard if it's, right? And so, so I think the way this plays out is, is as more and more of the models do more and more of these things, the marginal returns to investing in just like having a reasonably intelligent harness that can make sure you're using the right model for the right job in the right moment, uh, goes up a lot. And both for individual developers who, you know, they want a fast answer, they want a correct answer, and they don't want, you know, any necessary waste or even really like the cognitive overhead of having to decide every time you interact with a coding agent, oh, like what model should I use for this one? Harrison Chase: I, I was going to ask, do you let users of Devin choose what models they use? Russell Kaplan: For Devin desktop, which is, uh, what we rebranded WinSurf relatively recently, and for our CLI we do. People, developers like the individual control when they're working with local agents. Um, for our, for our cloud agent, we don't. It's a, you know, we'll have like Devin Fusion, which we talked about, which is the sort of, um, frontier performance with cost optimization option. We have like a more affordable agent, we have, you know, the like the Max or Ultra agent. So we have this like sense of kind of tiering, but I actually think it's like a UX bug in the fullness of time for people to have to think about this for their cloud agents. And, you know, we saw this early days we ran some tests where we allowed people pick the models. And then two things happen where, one is, sometimes people would then give a task to Devin and it wouldn't work and they would complain, oh, Devin was so dumb here. And we looked at that, well, why did you pick this model? But actually, that's kind of more our fault, I think, than the user's fault. And then the same thing would happen where they would run tests like, oh, this was like really expensive. Uh, well, why did you use this very expensive thing? And so, so I think that the natural equilibrium is people want the best performance for the best price without compromising on correctness, right? And so, folks who are building agents, I think have, in some sense like a user experience responsibility as well as, you know, just building good price performance products, to try to do as much of that optimization as possible. And that's what we recently put out, uh, with Devin Fusion, which is the sort of next generation of our own harness designed for frontier-level capabilities, but with maximum price performance. And so what we found is that, by being a little bit more clever about some of the routing and the decisions of like what you're doing with which model and, and thinking about this in a cash aware way, uh, we can get about 35% you know, better price performance with actually a slight, a slight increase in quality. Um, and I think that's sort of where a lot of this needs to go for the technology to be continued to be adopted at the at the rates being adopted. Harrison Chase: So talking about that like a little bit more, because there's there's routing and then there's also, OpenRouter launched OpenRouter Fusion, which is not routing, it's running it on multiple models in parallel and then combining things back together. So when you guys have Devin Fusion, is it, is it routing? Is it like running multiple things? Yeah, so it's, it's doing both. So, so we we put out a technical blog post that shares a little bit more detail on our implementation, um, but one core component of it is this idea of a sidekick agent. And so, you know, what we'll do is we'll have the sort of frontier quality model executing on the task, and then in parallel we'll have a more price-performance model executing on the task, and then there will be some decision making that the, the the frontier quality model has to do of, okay, well when do I delegate to my sidekick? But having them both work on the same task in parallel, it lets you make sure that there's still context for both of these agents, and we try to share a lot of context too, and we do, you know, writes to the file system frequently if there's if there's things that are exploding out of context, but having basically both work in parallel, then frontier model knowing, ah, okay, I can pass this off. And actually, as the models get smarter, the most frontier models get smarter and smarter, they're one of the key skills we see is they're like way better at delegating too. Like Fable is very good at knowing, ah, this task, I can delegate to a dumber model and it's going to be okay, which kind of makes sense if you think of human career progression also, like one of the aspects of of growing in your own career is learning how to delegate tasks and how to, like, do the highest leverage tasks yourself. And so we kind of see the same technical pattern in, in the models. Harrison Chase: You guys have your own set of models as well. SUI 1.6, I think is the most recent one. Why did you guys train that? What do you see people using it for? Russell Kaplan: Yeah, this is a great question. So sometimes people ask us, you know, why even RL or post-train your own models at all? Like aren't the next models from the frontier labs just going to be better and, uh, and better and better. And definitely, and we get super excited every time new frontier lab models come out that are better at, um, their task. But I think there's two important reasons for us to be spending a lot of energy on our own RL and post-training for our models. So, the first is that you actually can deliver frontier capabilities at any given moment in time through greater specialization. And I'll give you a recent example, which is, we shipped a product, uh, we shipped a product recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you know, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale. You know, when, when we tested their their their chips, we were really confused because we were like, this is really good, you know, like it's, it's, uh, it operates a unique point on the Pareto curve of price and throughput. And you know, we were getting about 950 tokens per second on our own models, which was many X faster than what we could get on, on GPUs for the same size of model. And it was a little bit more expensive for us to serve. They were great partners and they enabled us to ship a product, uh, we shipped a product, uh, relatively recently called Devin Review, which you guys are, uh, are using in really interesting ways inside, inside LangChain. Yes. To deliver Devin Review, we wanted to not just have sort of a human interface for understanding diffs, uh, and kind of grocking large amounts of AI-generated code quickly, but also to be able to run quick static analysis and lightweight, um, kind of model-driven analysis on, are there bugs in this code, are there security vulnerabilities in this code, are there things that we should be, sort of, linting and automatically checking. We can use, um, frontier models for that, uh, we can also use cheaper models. But then, uh, you get really tough trade-offs on the price performance curve, you know. You don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price-performant model for a very specialized task. And we know that that model has a half-life that's going to expire. And that's fine, because by the time it expires, we'll be very happy, and we'll be work on the next specialized model for the next workload. And so, I think a lot of kind of building an AI startup right now is, is being very willing to think in these like three to six month increments of, okay, given this state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization. That's part one. The the second reason we work on this is really helpful, is it does go back to price performance, which is, as more and more models are capable of doing more and more things, how can we continue to deliver, you, the best possible price performance for our customers, um, while it starts with, if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SUI 1.6 is actually the most popular model in Devin Desktop, by number of tokens consumed. It's about Opus 4.6 level, um, and we'll have more to say on on new models, uh, on new models coming out soon. There's going to continue to be this, like, frontier of capabilities that use the most, uh, frontier models for. And startups, not just us, I think more startups should be considering how do I get, how do I make my own models that are specialized for my domain, because more and more of the tasks in my domain are going to be doable by any model. Harrison Chase: How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones, or, yeah, order of magnitude? Russell Kaplan: It ranges. Yeah, I, I would say it's like single digits number of specialized models. Because because part of our own focus as a company is on software engineering, and so, you know, having like the SUI 1.6 series, that model series is something we can, we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models, but for the areas of our product that we think it's going to make the biggest impact, uh, not just on performance by the, or like kind of quality, it could be literally on like latency and speed and and, you know, affordability, we want to, we want to make sure we continue to invest there. And you guys optimize a lot for latency with SUI 1.6, right? Russell Kaplan: Yeah, one thing that was really fun, um, uh, work on SUI is, we were the first kind of western firm to deploy Cerebras at scale

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Artikel:The Cognition Team

    Devin Fusion: Multi-Modell-Architektur senkt Entwicklungskosten um bis zu 60 Prozent

    Cognition stellt Devin Fusion vor, ein System zur dynamischen Verteilung von Programmieraufgaben auf mehrere KI-Modelle. Durch die Kombination eines führenden Frontier-Modells mit einem günstigeren Sidekick-Agenten und dynamischem Routing während der Sitzung bleiben Qualität und Problemlösungskompetenz auf Spitzenniveau, während die Ausführungskosten um bis zu 60 Prozent sinken.

    KI & AI· Ankündigung

  • Artikel:The Cognition Team

    Produktivitätsmessung autonomer KI-Entwickler: Cognitions Ansatz für Devin

    Cognition stellt ein automatisiertes System vor, das den tatsächlichen Wert von KI-Programmiersitzungen in produktiven Entwicklerstunden beziffert. Da Token-Verbrauch und reine Codezeilen schlechte Indikatoren für echten Ertrag sind, analysiert ein separater Agent die Arbeitsprotokolle von Devin, filtert unproduktive Sitzungen heraus und schätzt den menschlichen Zeitaufwand konservativ ab.

    KI & AI· Forschung

  • Link:Cognition

    Cognition: Plattform und Entwicklungsstand des KI-Software-Entwicklers Devin

    Cognition präsentiert Devin als autonomen Software-Entwickler, der Programmcode eigenständig plant, schreibt, testet und bereitstellt. Das System arbeitet direkt in bestehenden Repositories und Toolchains von Unternehmen. Ziel ist es, menschliche Entwickler von Routineaufgaben zu entlasten, damit sie sich auf Systemarchitektur und Problemlösung konzentrieren können.

    KI & AI· Tool

  • Artikel:Rohan Choudhury, Carlo Baronio, Ben Pan, Sam Lee, Eric Lu, Steven Cao, Joe Li, Andrew Wang, Adam Zweiger, Ray Wang, Gary Chang, Silas Alberti

    SWE-1.6 von Cognition: Fokus auf Modell-UX und parallele Toolnutzung

    Cognition hat das Modell SWE-1.6 für Software-Engineering-Agenten in Windsurf veröffentlicht. Neben hoher Benchmark-Leistung setzt der Hersteller vor allem auf eine verbesserte Modell-UX, um Gedankenschleifen zu reduzieren und Tools parallel auszuführen.

    KI & AI· Ankündigung

  • Link:The Cognition Team

    FrontierCode: Neuer Benchmark für Merge-Fähigkeit von KI-Code

    Cognition hat mit FrontierCode einen Benchmark und ein Leaderboard veröffentlicht, das bewertet, ob von KI-Modellen erstellte Pull Requests von Open-Source-Maintainern tatsächlich gemergt werden würden. Die Bewertung stützt sich auf Unit-Tests, Verifizierer und praxisnahe Rubriken von Maintainern.

    KI & AI· Sammlung

  • Artikel:Scott Wu

    Cognition führt Produktivitätsgarantie für KI-Entwickler Devin ein

    Cognition bietet Unternehmenskunden eine finanzielle Garantie für den Einsatz des KI-Entwicklers Devin. Ein automatisierter Schätzer bewertet den tatsächlichen Zeitwert der gelieferten Entwicklungsarbeit; bleibt der Gegenwert hinter den Vertragskosten zurück, erstattet Cognition Guthaben von bis zu 10 Millionen US-Dollar.

    KI & AI· Ankündigung

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.