Wöchentliches KI-Update: Neue Modelle, Agent-Harnesses und Entwickler-Tools (August 2026)

VideoLev SelectorNews

Lev Selector fasst die wichtigsten KI-Entwicklungen von Mitte August 2026 zusammen. Schwerpunkte sind neue Modell-Releases (u. a. Muse Glimmer, Grok 4.6, DeepSeek V4 Pro, NeMo Tron 3.5 Lightning), modulare Agent-Architekturen wie Prime Agent und DeepSeek Harness sowie Werkzeuge zur Token- und Kostenoptimierung beim Agent-Coding.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Meta veröffentlicht Muse Glimmer (Apache 2.0, 30 Mrd. Parameter), ein multimodales Modell für lokale Always-on-Agenten mit 128k Kontextfenster und Apple-MLX-Unterstützung.
  2. xAI bringt Grok 4.6 mit 500k Token Kontextfenster heraus; das Modell soll Benchmark-Werte auf dem Niveau von GPT-5 6 Soul erreichen.
  3. Prime Intellect erzielt mit Prime Agent Harness laut Bericht 95,5 % im ARC-AGI-3-Benchmark durch rekursive Python-Arbeitsumgebungen statt vollständigem Prompt-Token-Laden.
  4. DeepSeek veröffentlicht DeepSeek V4 Pro zusammen mit dem quelloffenen DeepSeek Harness v0.1 und passt seine API-Preise an.
  5. Nvidia stellt NeMo Tron 3.5 Lightning vor: Ein Hybrid-Modell (30 Mrd. Parameter gesamt, 3 Mrd. aktiv) mit 52 Layern, davon nur 6 Attention-Layers sowie 23 Mamba-2- und 23 MoE-Layers für reduzierten GPU-Speicherbedarf.
  6. Claude Code erhält Cross-Session-Messaging; externe Utilities wie Caveman, Ponytail und RTK (Rust Token Killer) reduzieren den Tokenverbrauch bei Reasoning- und Shell-Ausgaben.

Warum das relevant ist

Der Fokus verschiebt sich spürbar von reinen Basismodellen hin zu spezialisierten Agent-Harnesses, hybriden Architekturen (wie Mamba 2) und Kostenreduzierung im laufenden Betrieb. Für Entwickler und Unternehmen wird der Scaffold – also das Laufzeit- und Kontextmanagement um das LLM herum – zum entscheidenden Hebel für Genauigkeit und Wirtschaftlichkeit.

Einordnung

Die gezeigten Updates verdeutlichen eine zweigleisige Entwicklung: Einerseits drängen extrem effiziente, speichersparende Architekturen auf den Markt (Nvidias Mamba-MoE-Hybrid oder Metas quantisiertes Muse Glimmer für lokale Hardware). Andererseits zeigt das Benchmark-Ergebnis von Prime Agent auf ARC-AGI 3, dass Leistungssteigerungen zunehmend über Werkzeugsteuerung, rekursive Variablenverwaltung und automatisierte Fehlerkorrektur (Self-Correction) erzielt werden, anstatt Kontextfenster unkontrolliert anwachsen zu lassen.

Transkript

Vollständiges Transkript anzeigen (4.341 Wörter)
Artificial intelligence updates every week on Friday 2:00 p.m. Eastern. Today is Friday, August 14th. We have a lot of updates. And the slogan for today is the first rule of autonomous systems, make recovery more automatic than the mistake. Oh, okay, fine. So, let's first look at leaderboards. Uh this was actually updated not 11th, but 12th. But generally, [snorts] you see the same picture. Uh I want you to see the Gemini 3.7 flash, which was just released. Right? Uh so, we have left is just a regular chat, and right is coding. And coding is heavily dominated by blue color, which is Claude. And Gemini is on the bottom of this, and OpenAI is yellow. Uh green is open source, usually coming from China. So, you see Kimiko 3 Max, Qwen 3.7 Max. Uh actually, we already have Qwen 3.8. Um anyway, so this doesn't change much. Uh next, this artificial analysis intelligence index. And what I want to point out that there are now the whole selection of very good models at a very cheap prices. So, we see GPT-5 6 Luna, uh which is the smallest in the GPT series, and it's very very affordable. It's the same level as uh DeepSeek, for example. Where Where is DeepSeek? This is DeepSeek. It's even less than DeepSeek. Uh Oh, no, this is Muse. Where is DeepSeek? Here. Oh, okay. The Oh, I know. The This is Pro. What What What What they did, starting this Sunday, they are raising their prices. Their prices are still like 10 times less than the American Frontier models because they really increase the quality recently, Deep Seek, and they also added the hardness, which is open source and free. So, they basically like you know, you're using Claude code and everybody likes this hardness, but it's closed code, right? You You cannot change it. But Deep Seek released the open one. Anyway, so you see this shows models executing the same tasks and the overall price, like price per intelligence cost per intelligence. You see on the right we have Claude Fable 5, very expensive at $3. GPT Soul, which is the top GPT model, is three times less. And Muse Spark 1.2, this is from Meta, is only 40 cents. And if you look here, Gemini 3.5 Flash Light, it's a good model, only 10 cents. You see, and Muse Glimmer just released, good model, only 7 cents and and so on. Okay. Here is information about Muse Glimmer. So, it was just released. And the word glimmer means a faint, unsteady, or intermittent light, like a candle flickering in darkness or distant lights barely visible at night. So, Muse Glimmer, flickering, sort of. And it's available via Ollama and many other. So, it's a Meta Apache 2 license open weight 30 billion multimodal for always-on local agents. So, it's a really good model for local work, text and image perception, tool use, multi-step long horizon tasks, and failure recovery. 128,000 contacts, which is not a million, but still pretty good. And it runs on Apple. You see the MLX. This means it has speculative decoding and image input, fast visual coding, uh launch Cloud Code, Pi Code X Open Claw, and Hermes through Alama and select different level. So, great. And there are variants, full precision. Now, it means 16-bit, and you may have heavily quantized, and then it's only 24 GB memory required. So, it fit in regular GPU. Anyway, next, Grok 4.6. It was released, and people are raving about it. It's It's It's really good. Uh so, it was released 2 days ago, uh available only to paid uh customers of Grok. And uh on the benchmarks, it's on the level with GPT-5 6 Soul, which is the top model from Open AI, which is frontier model, very close to Cloud models. 500,000 token context window. And [snorts] pricing is $2 and $6 for input output tokens per million tokens. Okay? So, this is for long-running agents. It's very good with visual tasks. Very, very convincing presentations on YouTube. I highly recommend to watch. Uh next, uh GLM. Uh previous version was 5.3, new version is uh previous was 5.2, now it's 5.3. And they also have a version which is specifically for coding, and you see how it behaves in comparison with other models, So, it's very good. 1 million token context, a sparse mixture of experts, same size as before, 743 billion parameters. Uh but what they were focusing on on better large-scale post-training real-world coding tasks, reinforcement learning environments and compute, uh frontier coding coding performance, creates plans, executed, and Price-wise, you see API pricing is quite affordable. 1.4 in and 4.4 dollars for out million tokens. They also have monthly subscriptions. Uh so, you can probably go with this light uh pro or max subscription. Um next, DeepSeek V4 Pro. Gosh, you see how many new models this week? So, the model itself and the DeepSeek Harness version 0.1. So, it's kind of preview, but it's open source. This is the first time they give their harness. And starting this Sunday, by the way, they increased the pricing. So, you see this is the new pricing. It's still much less, like order of magnitude less than American frontier models. Uh but look, for example, uh depending on peak off-peak hours, and this means Chinese peak off-peak hours. Uh but still, it's even the heaviest is less than $2 for flash, and for pro, it's less than $4 output. So, it's still very very moderate. Any Anyway, so this is high high-quality model, right? And apparently high-quality and open-source harness. It's a separate product. Uh Gemini 1.5 Flash was released. Uh Brin, co-founder of Google, he is not an executive, but as a co-founder he joined the development and he is pushing to move Gemini because there were delays in releasing the next version of Gemini Pro. But they they just released flash and it looks very good on on benchmarks and it's quite affordable on the same level with Chinese models. So and it's available through different Google channels. Next, OpenAI ships CPT 5.6 cyber. So it's a version of their model for cyber security, advanced cyber security work. Cloud code cross-session messaging. So now when you have multiple sessions, they can talk to each other. Exchange concise plain text handoffs reducing manual copy-paste. So this is a new feature very important. Bytedance is secretly building 10 trillion parameter model, but it's in training so it will take several months until they will release it. This is plug for my channel so more than 7,000 subscribers and 300 videos. Started doing also shorts automatically generated almost. So name of the channel Left Selector, please subscribe. You can download slides. I provide the links under the video for GitHub and for Google Drive. And please pause the video and answer pinned questions under the video or just any comments, any feedback welcome. Next, Discovery Loop. It's a new company which was formed by some of the very famous people from Google starting with Jeff Dean who is absolutely legendary, and he took these people quickly. I remember Google Translate when the times and like magazines, he was the like central person. All these people are very very famous. So, it's like a star team, and they left Google. And they're actually sponsored by Google money-wise. And their company is called Discovery Loop, and their goal is to create recursive self-improvement uh system, right? Why is half of modern Google? Yeah. Uh in the because nothing Okay, good. So, we'll see. Continuous exploration, so it's a model for research, for science, for improvement, self-improvement. Uh next, Grok uh imagine image 2.0. So, this is latest version of the imaging processing. Uh next, Hermes browser use upgrade. So, you know, Hermes is a very famous, very popular uh agentic framework, open source. And now they kind of consolidated 12 per action browser tools into one browser use mode and CLI. And it uses less tokens and does better job. So, good upgrade. Uh next, Grok bot agent team. So, this is only available for people who paying money. It isn't better. And it is good, but uh you don't receive a lot of credits. So, after like 15-20 minutes, you're out of credits. Uh but, yeah. Uh next, NVIDIA NeMo Tron 3.5 lightning. This is very very interesting. So, this is model uh which uh like here these numbers. It has 52 total layers and only six of them are attention layers. 23 are Mamba 2 states space layers and 23 mixture of expert layers. So, not a lot of attention and because of that it doesn't require as much memory. So, this means if you running it in production, you can run more parallel tasks. So, it's much more practical. It's It's not big. It's only 30 billion parameters, 3B active and it can run on a single GPU. >> [snorts] >> So, So, again, it's mixture of experts and it's Mamba 2. Mamba 2 is this state space layers. Uh less GPU memory required, max context length 1 million. Well, depends on GPU and memory. Uh provides inference focused 4-bit uh quantization and reference fine-tuning and also provide the 16-bit uh okay by Hugging Face and Nvidia. Tool use, structured output, code validation, sub-agents, multi-talking predictions and and so on. So, yeah, this this is a good model. This is a very good model. Um Anthropic is adding invisible watermarks to Claude-generated content. Uh and they promised that they will also provide tooling to see if the text contains the watermark. Okay. Claude Obsidian plugin v.2. So, uh I guess you can use it to generate Obsidian knowledge base. I was doing it without plugin, just giving instructions to Claude and it was doing pretty good job, but here is the plugin you can use. A gold card seed exploit. So, what happened? There is a hardware device for Bitcoin. It's It's a wallet. It's It's a cryptocurrency wallet. And it had a problem with seeding its random generator. And because of that, people were able to steal 100 million in Bitcoin. >> [snorts] >> And I guess it was fixed. I don't know how they came up. It's probably using AI. But anyway, [snorts] this is a big event. Um Prime Agent Harness versus Hermes. So, Hermes is an agent, so it is a harness. A Prime [snorts] Agent it's also a harness. It's a from a company called Prime Intellect. Uh so, it reports 95.5% on Arc AGI 3. It's difficult to believe because just several months ago, people were struggling with single percent. Now, it's 95.5. And for humans, it's 95.4. So, it's basically better than humans. Absolutely amazing. Agent performance depends heavily on the scaffold, context management, tool design, memory verification, retrieval loops, and so on. So, it's MIT license, so it's permissive open-source license. Uh Open Agent Harness. Uh wraps existing LLMs with persistent Python environment tools and subagents. Recursive language model approach, keep large input external as variables. Uh the model writes code to inspect, filter, summarize, and delegate work instead of ingesting every token. Uh yeah, this is amazing amazing achievement. This week is like crazy. Uh long context models suffer from context growth, accuracy declining. And this model kind of defeats that. Recursive language models inspect stored document slices instead of loading everything in one prompt. Okay, so this is Prime Intellect again and you you see primeintellect.ai. Uh Prime Agent versus Hermes. So Prime Agent excels at deep long-running coding and research using persistent Python sessions. Hermes excels in connected agents tools live search and CP. So this like a day-to-day practical work. Next, uh >> [snorts] >> uh Gauntlet Loop AI QA. So Gauntlet is this thing. It's a protective kind of Uh so this is by Matt Shumer. He's an entrepreneur and investor. He is famous for this essay "Something Big Is Happening" uh where he was saying that public is underestimating rapidly proving AI capabilities. And you see 80 million views and there was a lot of discussion. Um so now he says that you can run the Gauntlet, meaning endure difficult series of attacks, tests, and criticism. So meaning that when you AI is doing something for you, then you can run it through a lot of tests and then you can return and improve, return and improve. So the prompt for the Gauntlet Loop needs three essentials. Ambitious outcome, concrete benchmark critics inspect, and explicit stopping boundary. So it goes in the loop improving the response. Uh [snorts] I guess there's nothing new, but this actually really really improves the response. Okay, next. Uh Harness R1, outcome graded agents. Deploy AI agent and collect every failure trajectory. Separate a separate 9 billion parameter hardness engineer reads those failures. Diagnosis keeps what's going wrong. Engineer writes executable code patches that change how runtime. When I say engineer, this is actually AI. >> [snorts] >> And validates actions and triggers recovery. Patched agents rerun the exact same task with original weights frozen. Only patches that actually raise task success survive. So it's kind of evolutionary approach to improve the hardness. The hardness can add start of task guidance decision hands and so on. It raises vanilla target average success from 44% to 50 54%. So coalition outcome tested hardness engineering can outperform much larger prompted generated models. So you reward the hardness for success. It's a new way of training. Wipe coding startup lovable. So these are the founders. Fabian Hedin and Anton Osika. They raised another 400 million at 13.3 valuation. So lovable hit half a billion annualized revenue return in June just recently. So this company is in Europe. It's a wipe coding startup. Okay, Mark Zuckerberg. So he discusses the future. He claims that well people should all have access to AI. It should be democratic. But at the same time people who have more money can get more AI. This is absolutely obvious. So it's kind of contradiction. Free or affordable AI, open models, individual agency, invention, entrepreneurship, tutoring, and so on. Uh yeah, so this is a discussion uh what will happen with AI and how to avoid concentration of power. So, he definitely against uh just few companies controlling AI. Double descent explained. So, this is this is interesting. So, when you in regular machine learning, when you're training the model and at first the error goes down, down, down, but then uh it starts going up again because what what's happening uh you're starting to overfit. So, um there is a uh uh trade-off and you want to find a sweet spot. Now, but when you're training a large model, not not a regular machine learning, but deep learning, large model, what's happening at first again, you increase it, then you start overfitting, you're increasing and at this point you have more parameters in your model than you have data, right? And after that suddenly it start decreasing again. And this is called double descent. >> [snorts] >> And you see push past the threshold and test error descends to a second time, often below the classical sweet spot. Once there are more parameters than points, infinitely many project perfect fits exist. And gradient descent quietly prefers the smoothest, lowest norm one. >> [snorts] >> This is double descent. So, in in interesting. Okay, um Simux. I was introduced to this application. It's a terminal for Mac and it is very convenient to run different agents. So, in the same application you can run cloud code and code X and what whatever you want and uh very impressive. I was switching from Visual Studio Code to Zed and now I'm really thinking to switch to this C Mox. And these are the links you can easily download, install, and try it. Okay. Yeah, this is interesting. This is a young woman, an influencer, and she created Vibe Coded application for the phone and she made a lot of money very, very quickly. And of course, she already had a lot of following, but still it's very impressive what you can do with AI. Uh Vid Fabric AI video test. So, Julie McCoy tested vid.io. So, you take the picture, it may be picture of yourself, you you take voice, and the system generates the video where you have lip syncing like avatar, everything looks pretty good. Uh for short videos. You see 10 to 30 second outputs. Okay. Stanford Virtual Biotech. Again, this is very interesting project. Um what they do, they use thousands, actually tens of thousands. You see 37,000 clinical trial agents. Process 60,000 trials. And and it actually works. So, this is for science, this is for pharmaceutical testing, quite an achievement. So, it's a Stanford Virtual Biotech and this person is James Zou, who is a researcher and who did this project. Okay. AI dev platforms, GitHub versus Vercel versus Replit. Well, of course, you've heard about all three of them and they have the pluses and minuses. And I definitely use GitHub and I do not use Vercel or Repl it, but I know people who use them very successfully. Okay, next. Uh Auto Magical Life. Uh so, Peter Diamandis uh wrote an essay. He fantasizes what will happen in just few years, like actually in '28, which is 2 years from now. MBTI will continuously use uh consented personal data to anticipate needs and act without explicit prompts. It would combine calendars, conversations, home cameras, location, wearables, glucose data, health records to optimize decisions in their environments. So, he predicts that AI will enter our daily life. Uh Granian versus Unicorn for uh web applications. Uh Uvicorn is a very common uh server. So, when you're creating a fast a fast API uh uh ap- application. Uh >> [sighs] >> Okay, so this is the app and this is sub process run. And here here you use Granian. Okay, so th- this is code it it shows how easy it is. It It's basically the same what you what you're doing with Uvicorn. Uh So, when when you're doing a fast API application, the difference though is that Granian uh is is written in Rust and it supports HTTP 2 out of the box. So, it's a synchronous with all all the latest protocols supported. Uh next. So, i- i- if you're making a new application, you probably should go for Granian and not for Uvicorn. Uh LangGraph runtime over diagrams. I tried LangChain LangGraph it 2 years ago. I was very upset because it's a complicated uh a system where documentation is always obsolete and uh where they constantly update things moving very fast and things break and you have no control over it. But uh people still using it. Uh there are a lot of downloads. It provide convenient environment. And uh maybe today uh when AI can look at your LangGraph application and fix it and figure out what's going wrong, may- maybe it it is uh useful. Uh they it's definitely production-grade, production-ready. And uh they put a lot of effort like for example, if something goes wrong and you need to restart something, rerun, uh it's all possible. Okay. Uh AI coding skills, modular control. >> [snorts] >> So here uh this person uh Matt Pocock, uh he uh teaches uh that uh you use modular approach like for example, Groom me interviews users uh one decision at a time. Two Spec preserves the agreed design in a code-free specification encouraging agents to inspect. Uh Two Ticket slices work by the end-to-end features rather than database API. Implement emphasizes uh test-driven development. So we are talking about skills which help you develop uh better applications. And uh yes, okay. Uh next. Karpathy prompting 2.0. This is unusual. So you know Andrej Karpathy is very famous. What what he found that you can start talking using your voice, not uh writing prompts, but talking to AI and kind of expressing everything you think about the project, right? The intent, uh the architectural decisions, whatever. So, start by warning the model that speech recognition may introduce typos, then freely describe goals, constraints, uncertainties, examples, and so on and so on. The model can reconstruct that messy content into a cleaner shared understanding, sometimes through a short interview-style back and forth. So, at the beginning of the project, he recommends this like as a starting point. And it's called Prompting 2.0. Okay. Next, Sand X. Harness influence, local AI performance. What he's talking about that you can take the same model, but use it with different harnesses, and you can see drastically different accuracy, different level of performance. OMP, Oh My Pi, terminal-based AI coding agent. Okay. Next, create Blender animations, AI and MCP. I recently see a lot of shots which are made in Blender. So, for example, a short video shows how jet engine works, turbine, whatever. It's animated video. And what you can do now, Claude code, you can use MCP to instruct Blender what to do, and you can create animations with text prompts. And it's actually becoming very, very easy. So, you see blender.org, it's a very famous tool for graphical design and animations. And now you can control it using text prompts and generate those animations and generate videos. A lot of demonstrations on YouTube. Uh Minimax Code 2.0 was released. Rebuilt desktop agent on the open source Pi framework and over 90% lower first token latency, so much faster, fewer stalls, stronger persistence. I should have put it in the beginning of presentation. It's actually very important. The teams of agents can divide the search strategy, production, and review among separate agents. Memory and reusable skills preserve organizational Okay, so this is very good development and it's basically a desktop agent. Uh Link 30 tiny, small, sparse, and fast. It's a hybrid reasoning mixture of expert model with less than 8 billion total parameters. 128 experts. Wow. Uh 266K contact window and different precision of the weights. Different levels of quantization. So this is good model. A Claude code token leak fixes. Uh what we're talking about? Why it's a leak? Because when a Claude is thinking and it prints text, you see the reasoning. It consumes a lot of tokens. And there are multiple ways and there are multiple tools which are specifically designed to reduce it. So here, for example, normal authentication likely fails because the token expiry check uses local time. A caveman. Caveman is one of these four tools. Auth fails local time expiry check, use UTC. So you see it's it's shorter. Ponytail does similar thing. Omniroute allows you to select different models and providers. 268 providers. Uh RTK, it's a rust token killer. Reduces shell output and so on. Uh so you can use these tools, one of them or in combination to reduce uh the tokens you're using and pay less money, I I guess. So, these are all the links you need. Uh, Kestra governed AI agents. Kestra combines deterministic orchestration with AI reasoning. Uh, keeping high-impact actions under human control. YAML close, versionable, reviewable, deployable, and GitHub uh, triage, post a human approval, can self-host locally with Docker and Kubernetes. Uh, I guess we can consider that. Uh, Hermes {slash} learn. Uh, so, what you can do, suppose you have a lot of documents and you can issue a learn command and it will convert those different documents like PDF, uh, book, the document collection into agent skill. So, you can create skills from documents and it's called learn. Resulting skill can organize source material into focused on demand references and so on. So, this is in Hermes, but you can create similar skill easily. What's going on? Sender, please mute yourself. Okay. Pax Silica. So, this is a big project and it started recently and you see there are many, many countries participating and they started in Philippines actually, the first place. Uh, so, it's US led initiative for secure AI semiconductors, minerals, energy, like everything. And they now have this place with more than 1,000 hectares for 4,000 acres, right? And so, they they're building stuff. It's all for AI. Uh open mouse bot agent teams uh your own team of AI bots in a chat application. Free open source mock agent workspace by this developer modeled on Grok bot presents a agent that chat contacts each with distinct role and they cooperate and work together and it's open source so you can use it. Okay. Sybase. I spent years working in Sybase and Sybase is very interesting history. Um was created by very talented people then 4 years after created they made a contract with Microsoft. And after several years Microsoft didn't renew the contract so Microsoft was supposed only to sell Sybase database. But when they were not renewing they offered to pay Sybase uh for their code and Sybase sold them the code and what became Microsoft SQL Server it was originally Sybase. And because of that Sybase uh well, it was not killed completely but you see in 2010 it was sold to SAP for almost 6 billion dollars. You see this Sybase and the product which used to be Sybase Adaptive Server uh Adaptive Server Enterprise uh it became SAP Sybase Adaptive Server Enterprise and IQ. Oh, they invented IQ. IQ is a database for analytics where it is indexed on a column by column basis, not row by row but column by column which is better for analytics. Uh Sybase is amazing company and they contributed so much and I I love them. This the simplicity, the elegance uh amazing. Okay, this is uh about jobs and layoffs. Nothing major happened in August yet. This is me as usual and thank you.

Links und Tools aus diesem Beitrag

94 weitere anzeigen

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Video

    Video:Lev Selector

    Wöchentliche KI-Updates: Modell-Releases, Prompt-Tricks und Agenten-Architekturen

    Lev Selector gibt in seinem wöchentlichen Rückblick vom 10. Juli 2026 eine Übersicht über aktuelle KI-Entwicklungen. Die Themen reichen von neuen Modellversionen (OpenAI GPT-5.6, xAI Grok, Anthropic Fable 5) über kostensparende Coding-Workflows und Prompt-Techniken für CLAUDE.md bis hin zu Fortschritten bei Spekulativem Decoding und Open-Source-Inferenz.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliche KI-Updates von Lev Selector (7. August 2026)

    In seinem Wochenrückblick fasst Lev Selector aktuelle Entwicklungen im KI-Ökosystem zusammen. Im Mittelpunkt stehen neue Coding- und Agenten-Modelle wie OpenAIs Astra und Metas Muse-Serie, Browser-Automatisierung sowie Open-Source-Fortschritte bei lokalem Speicher und Hardware-Kompatibilität.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: DeepSeek Harness, Model-Routing und Scaffolding-Trends

    Im wöchentlichen Überblick analysiert Lev Selector den Wandel der KI-Entwicklung: Der Schwerpunkt verlagert sich von immer größeren Modellen hin zu modularer Scaffolding- und Harness-Infrastruktur. Zu den wichtigsten Entwicklungen zählen der virale Erfolg des modularen DeepSeek-Harness, die Übernahme von OpenRouter durch Stripe sowie neue Modell- und Routing-Strategien.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Open-Source-Modelle holen auf, Agenten-Harnesses und Metas KI-Ernüchterung

    Lev Selector fasst in seinem wöchentlichen Rückblick aktuelle Entwicklungen der KI-Branche zusammen. Im Mittelpunkt stehen der rasante Vormarsch offener chinesischer Modelle auf Plattformen wie Vercel, neue Benchmarks durch den Nvidia AVO Agenten bei ARC-AGI 3 sowie die explosionsartige Verbreitung von Open-Source-Coding-Harnesses. Zudem beleuchtet der Vortrag Metas Erfahrungen nach Entlassungen, bei denen unzureichend überwachte KI-Agenten zu einem starken Anstieg von Vorfällen führten.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Neue Modelle, Agenten-Workflows und Infrastruktur

    Lev Selector fasst die wichtigsten KI-Entwicklungen der Woche vom 24. Juli 2026 zusammen. Zu den Schwerpunkten gehören aktuelle Modelle von Anthropic, Alibaba und Poolside, Kostenoptimierung bei API-Aufrufen, neue Agenten-Funktionen wie Claudes Skill-Recording, Google-Sucherweiterungen sowie der Umstieg europäischer Behörden von Windows auf Linux.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Regulatorische Bremsen, GLM 5.2 und Harness Engineering

    Lev Selector fasst die Entwicklungen der KI-Szene Ende Juni 2026 zusammen. Zu den Schwerpunkten gehören staatliche Verzögerungen bei US-Spitzenmodellen, der Aufstieg von Open-Source-Modellen wie GLM 5.2, Fortschritte im Multi-Agenten-Orchestrieren sowie die wachsende Bedeutung von Harness Engineering.

    KI & AI· News

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.