Wöchentliches KI-Update: GPT 5.6, Claude Sonnet 5 und Ornith

VideoLev SelectorNews

In seinem wöchentlichen Überblick für Anfang Juli 2026 bespricht Lev Selector aktuelle Entwicklungen in der KI-Landschaft. Im Fokus stehen neue Modellveröffentlichungen von OpenAI und Anthropic, staatliche Zugangsregulierungen zu Spitzenmodellen, Framework-Kritik an LangChain sowie lokale KI-Tools und Infrastruktur-Automatisierung mit Pulumi.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. OpenAI GPT 5.6 wurde in drei Varianten (Sol, Terra, Luna) vorgestellt, ist aber aufgrund staatlicher Prüffristen zunächst nur Behörden und Partnern zugänglich.
  2. Anthropic hat Claude Sonnet 5 mit 1 Million Token Kontextfenster sowie Claude Fable zurückgebracht, setzt jedoch nach einer einwöchigen Promotion strengere Quoten und Sicherheitsfilter ein.
  3. Das kalifornische Startup Ornith trainiert Problemlösungsstrukturen direkt in Modellgewichte von Gemma 4 und Qwen 3.5 ein, wodurch kleinere Modelle teils größere Frontiersysteme übertreffen.
  4. DeepSeek hat Dispark vorgestellt, ein Werkzeug für spekulatives Decoding, das die Token-Generierung um bis zu 85 % beschleunigen soll.
  5. Kritik an LangChain: Hohe Abstraktionsschichten, häufige Breaking Changes und Debugging-Probleme treiben Entwickler zu schlankeren Ansätzen oder Deep Agents.
  6. Für Infrastructure as Code (IaC) vergleicht der Sprecher Terraform mit Pulumi und demonstriert, wie Claude Code IaC-Code in Python automatisiert erstellen kann.

Warum das relevant ist

Die KI-Landschaft zeigt eine zunehmende Kluft zwischen staatlich regulierten, geschlossenen Spitzenmodellen und immer leistungsfähigeren Open-Source-Alternativen. Für Entwickler und Teams werden pragmatische Workflows rund um Agenten, lokale Modelle und transparente Abstraktionen gegenüber überladenen Frameworks entscheidend.

Einordnung

Lev Selector skizziert eine Marktsituation, in der führende US-Labore wie Anthropic und OpenAI einerseits eng mit staatlichen Akteuren kooperieren und restriktivere Richtlinien durchsetzen, während Open-Source-Initiativen und Techniken wie Fine-Tuning mit Reinforcement Learning (Ornith) oder spekulatives Decoding (Dispark) die Effizienz massiv steigern. Bemerkenswert ist auch der Trend hin zu code-basierten IaC-Ansätzen via LLM anstelle proprietärer DSLs.

Transkript

Vollständiges Transkript anzeigen (4.015 Wörter)
Artificial intelligence updates July 3rd 26. We do updates every Friday at 2:00 p.m. Eastern time New York time. So, today is actually official holiday because tomorrow July 4th is Independence Day and not just Independence Day. It's exactly 250 years since in 1776 50 something people joined together and on this on July 4th they actually approved the text of Declaration of Independence and then later they voted. Yeah, but it's a big holiday, big event. And the epigraph for today is competition is the only solution to problems like government involvement and monopoly of Anthropic. Anthropic from small company become has become a real monopoly. Okay, let's move on. Hold on a second. So, these are leaderboards from 2 days ago July 1st. And as usual you see for coding it's all dominated by Anthropic by Claude models. Here are the sizes of all the models. You see that Anthropic actually they're very big models in trillions. All these models are big. Gemini doesn't look very competitive. Quen 3.6 looks good, right? And JLM 51 and now JLM 52. It is here on this leaderboard, but everybody is [snorts] very happy with JLM because they're open source. And you can run quantized version on your at your at your home. Okay. OpenAI GPT 5.6 was released and not released because government has now the right of first night or other first month. They don't allow GPT to just Open AI to give the model to the public and they want to keep it and evaluate it and let their government and some close partners use it. So there are three models Soul like Sun, Terra and Luna. Right? And they pretty good like really really good. Soul of course is the flagship highest capacity a Terra balance similar to GPT 5.5 and Luna is as usual fastest. So in Anthropic world this will be a Claude a Sonnet a Terra Haiku is a small one and Opus is the big one which is now a Fable. Now Fable is actually returning. It returned they started to roll it back on July 1st and now people have it and for the first week they give promotion. So for example if you have subscription Pro Max or like $100 $200 you can use all your subscription on Fable. But after July 7th you [snorts] can only use up to 50% of your quota and the rest it automatically downgrades probably to Opus. Well you can always pay using API token prices which are high which is like $10 or $50 input output per million tokens. So Fable is returned has returned but also they added more like governance security so it refuses to do many things simply because it suspects you that you're doing something illegal. Uh so it's it's okay. Uh now they released Claude Sonnet 5. You see this number five. It is cheaper and for the first 2 months it's even more cheap. So it's $2 for input tokens and 10 for output. And after that it will be three It's not 43, it's $3.15. And compared to Opus it's 5.25 and compared to Fable which is 10 and 50. Right? So this is quite affordable and pretty good model. But some people say it's actually more expensive than Opus because they updated tokenizer and it spends more on tokens. It has 1 million context window. Okay, next. Uh yes, this channel I love him. He said, "Well, why we have this one week to use Fable, let's do something interesting, take advantage of this." So he made several projects and he describes them. Like for example, you can replicate some paid software or cloud services as your own custom perform complete teardown analysis of how to use Claude code and make your usage more optimized, building agentic OS which is a custom web app wrapper on top of Claude code. Uh code review, deep debugging, complex long horizon software, and so on. Yes, Fable is very capable and now it's a good time to learn to use it. Nano Banana released version two light. It is faster, cheaper image generator. Now Google came up with Gemini Spark for Mac. Uh remember that Apple decided to use Google model Gemini as their main AI. And now we have this application Gemini Spark on Apple computers. Uh next Anthropic will require ID. Well, these were rumors. Um the problem is that uh there was a suspicion actually not suspicion, it's a proven fact that there were 25,000 fraudulent accounts which were bombarding uh Anthropic models asking questions, getting the answers, and then use this question-answer pairs to train their own Chinese models. So, Anthropic was thinking maybe to create something like know your customer a system where um they need to know the person. Well, if you subscribe and you pay money, of course you know that you're a person, but uh if you're using free access, uh just email, then it's not enough. They were thinking maybe they will require people to provide some sort of ID, but it didn't happen yet. Now, Claude Mython Mythus was also returned released to 100 more than 100 partners of government, and then they expanded it even more, but it's not open to general public. It never was. Uh Nvidia and Quen. So, Nvidia took Quen model 3.6, and they created their own version which runs as floating point 4-bit, and it's fast, it's good, it can run in a very small amount of memory, so basically on your laptop or computer with one Nvidia GPU. Uh 256K context length. Well, respectable, not million, but good. It's multimodal, understands text, image, and video. Uh It's interesting what's happening with these models. They became very capable for local usage. Um I was thinking about Wikipedia. Uh you can download the whole Wikipedia in one language, let's say English, and only latest versions of all the articles and compressed images. It would be like about 50 GB. 50 GB is not much. Uh my laptop has like terabyte, I think. Or maybe 500 GB. Anyway, anybody can now have all knowledge, all like big chunk of human knowledge on your laptop, and then use local model, because you don't need this model to reason, to be very smart. Uh it's like a librarian. All it needs to do is to find relevant information, and uh you can arrange this as some sort of wiki sto- Actually, it is wiki. Wikipedia, it is wiki. So, uh it may be very efficient. Uh knowledge at your fingertips, science at your fingertips, on your laptop, local. Hermes, uh mixture of agents. So, this is a popular uh thing then uh you use several models in parallel, and then go for consensus, right? So, uh many agents, and Hermes, of course, is one of the famous one, very popular ones, implement this technology. What it allows you to do, better accuracy, less hallucinations, better quality of responses. And uh yeah, you you see, you select mixture of agents preset like any other model, like {slash} model {slash} mixture of agents. Okay. Uh a good way to run Hermes, if you want, on Hostinger, uh it's called Hermes agent there. It's a one-click Docker VPS application. You can install it and run it there. Okay. Uh this video, I actually enjoyed it. Um What he's talking about that we see that government takes over. Well, it's not how it was with nuclear power where both America and Russia they just made it a secret military projects. But government interferes more and more. You see getting access to frontier models, creating widening capability gap between internal up technology and public tools. So now you have a split into two classes. You have government, government agencies and their close partners like a big companies and everybody else. And over everybody else doesn't have access to the latest models, latest tools and so on. Well, Chinese don't do that and they put their best models in open source. And now more and more we see that American companies are starting using Chinese models. US may label foreign models as radioactive security risk due to sleeper circuits. Well, it hasn't Well, no, they included Alibaba for example into blacklist but not deep seek and not GLM yet. So dark ages, two classes what I said government and big companies and others receive models later. Okay, or if ever because Mythus was released in April and it's still not available for for public. Anthropic's new Claude tag feature. So you can call Anthropic, you can use it from Slack and you can talk to it using at Claude. Like you say at Lev, at John, now you can do at Claude and you can talk to Claude from Slack. And this is a new way where your agent really becomes a team member. So, it can read Slack. It It can have this knowledge about company culture, communication, what to do, uh act as proactive, persistent team member, learns about your company team operations. Uh well, there are plusses and minuses. And minus that at some point the whole company will be different agents, not people. It's like a Trojan horse. Uh renting a digital employee from Anthropic. So, the danger here, imagine that you have a company which uses Claude as a team member. And then you have another company and another company. And then you have pretty much all corporation in America uh using Claude. More than that, Claude is a team member. Claude is a main uh software coder. Claude is a main manager. Claude is basically running the businesses. Claude is running the whole Am- >> [laughter] >> whole America business, all companies. Gosh. Kwen Agent World. So, this is open source from Alibaba. It builds a virtual world in inside AI to simulate agentic environment. Okay. Uh Open Tag. Oh, yeah. So, I I spoke about Anthropic Tag, but there is another project, Open Tag, which is similar. So, it works with Slack, WhatsApp, and others. It can use any model, not just Claude. And it can do the same thing. Right? So, it has its own copilot. Okay. Next. Okay, this is just plug for my channel. I reached 7,000 subscribers finally at 294 videos. The name of the channel Left Selector. I provide links to download slides with all these links. And also, please answer questions, provide comments under the video. Uh next. Uh OpenAI working with Broadcom uh to create chip, which is called Jalapeno. It's not done yet. It's in testing, although they planning initial deployments in data centers by the end of this year. Uh, Google Gemini SQL 2. So, it's a text to SQL. And it is new. It is very good. Translates natural language into executable schema aware SQL queries. Okay? Uh, Ornith. Oh, this is this is beautiful. This is the big big big thing. Um, so when you working with an agent, uh, there are two pieces involved. There are the model and the harness. And the harness is how to solve the problem. Uh, so maybe in school when teacher was teaching you how to solve math problems, it was teaching you a step-by-step approach how to solve mathematical problem or geometrical problem, how to think about it. So, here they took the same thing. So, you have harness and you have model and they trained uh, harness into the model itself, into the weights of the model. And they did it by kind of fine-tuning by uh, using reinforcement learning as part of the training to teach the model to use this harness. And the results are absolutely amazing. So, they created they trained several models of different sizes, 9 billion, 31 billion, 35, and almost 400 billion. And um, they built on top of Gemma 4 and Gwen 3.5, which are open-source models. And uh, so it's basically kind of reinforcement learning fine-tuning of the original models. The result is that their models are beating frontier models. Really. Or very, very close. You see here it says 9B beats 31B. Example comparing Gemma 4 and Ornith. Uh, the name Ornith comes from the ancient Greek word for bird. Right? And this is one of the four co-founders of the company. The company is in California in Santa Clara. And uh, led by Juwei Li. Has few dozen employees. Got 100 million in funding. This is absolutely amazing project. Uh, outstanding results. Uh, next, uh, Deep Seek Dispark. Uh, so it's a open source tool and it accelerates uh, the model. Uh, uses speculative decoding where small fast draft models predict a block of tokens all at once. Uh, so already used in production for Deep Seek V4. Delivers tokens up to 85% faster and can be attached to other open source models like when or Gemma. Uh, so this this this is excellent. Uh, Google open knowledge format, OKF. So this is an example of how it looks. It kind of like uh, how skills and Anthropic skills, but a little bit different. It has predefined uh, field names. And you can take your knowledge and you can express it in this format. And then once you have those files, this OKF, which is open knowledge format. Uh, once you have these files, you can make a wiki out of them and use it as your memory. And this is very good, very efficient way uh, to to build memory for your agent. Uh Google knowledge cat catalog ingest okay of bundles and expose them as a searchable catalog, but the format itself is a platform agnostic. Right? You can store the files in Git object storage or any file system. So, this is a good way to go, definitely. Okay, six power phrases for Claude code. I love this video. And we're actually using these phrases, but he formalized it. So, launch sub-agents. You tell your agent to use sub-agents for specific tasks. Write me an implementation spec. I always do that. I start with this specs. Interview me about the project. Verify before you build. Based on this conversation, build me a skill. So, it looks at the history of the conversation and builds a skill. Automate this. So, pieces of your work can be automated and you can explicitly ask it to automate it. The most common mistake is failing to provide enough context or failing to create a plan. Right. Now, ponytail. You see this ponytail here? I don't know. I don't like this logo picture, but that's what they use. Uh so, it prevents verbose behavior. Uh you see Opus 4.8 gets 71% faster, fewer lines of code, 53% reduction. And all it is, it is uh basically rules. W- W- When you're using Claude code, you provide rules how to format your software, how to approach it. And here there are rules, like a strict decision letter letter, lazy senior development mode. So, the agent writes only necessary code, avoiding verbose over-engineering outputs. So, you ask uh your agent to be concise. And this project it like put it in writing and you can just use their their rules. Uh Google Pommely AI marketing agent. So, this is the picture. Helps small medium-sized businesses create brand content, analyzes your website, uploaded materials to build business DNA profile capturing tone, colors, font, uh imagery style, generates campaign ideas, social posts, and so on. So, this is a marketing a Hermes agent OS a self-hosted always-on AI operating system built around Hermes agent, which is MIT licensed, which is completely open source. Autonomous agent very popular already. Think I spoke about it earlier. Uh it has multiple front ends like CLI in the browser, Telegram, Slack, Discord, WhatsApp. Model agnostic can use many different models. Uh supports MCP APIs, uh runs on different platforms. And now it has Jarvis, a voice-driven Jarvis style. So, you can talk to it and it's really like an Iron Man. Like it's really funny how it talks. Okay. Um CPT Cyber. So, these are just OpenAI models from OpenAI, which are fine-tuned for cybersecurity. They build on top of today's models, GPT 54, GPT 55, and used for discovering vulnerability, malware, and so on. So, it's a good like family of models specifically for cybersecurity. Uh law LangChain is dying. I had experience with LangChain 2 years ago, a little bit more than that, and I hated it. On LangChain provides library. They started with OpenAI models and how to connect to vector databases, how to run the models, how to do this, how to do that. But then they added more tools, more models, and it became more and more levels of abstraction, and it became very difficult to work with it. Uh you see heavy frameworks introduce complex abstractions, increase token costs, obscure prompts. Uh emerging Okay, well, the core takeaway is to always use smallest, simplest tool that successfully completes the job. Yes, LangChain is not the smallest tool. Uh what you've got is an agent inside an executor inside a chain. So, you really have like five, six, or maybe eight levels of abstractions. Uh your real prompt buried eight layers down under code you never wrote. So, developers are walking away. The abstractions are over complicated. The docs never match the actual working code. Uh you see the bug isn't in your 50 lines. It's somewhere in the framework of 50,000 lines. So, you end up stepping through abstractions you didn't write in a stack trace that most somebody else's code, and the ground keeps moving under you. LangChain rocketed past 100,000 GitHub stars and became just as famous for breaking changes between versions. The interface you learned last quarter quarter is deprecated. Popularity was never the same thing as stability. Yes, this was my impression. I I just tried to use it and then I jumped out and I never looked back. I don't want to use LangChain. Well, but this is me personally, but this looks like many people have the same experience. Oh, LangChain deep agents. LangChain introduced the deep agents and these are open source it's a agent harness to build agents. Of course, based on the uh technology. Includes plumbing like systems, context management, sub-agents, planning, and so on. Developers start with complete functioning terminal agent. Okay, built on LangGraph deep agents code and model agnostic and so on. Well, LangChain is very famous despite that I don't like it, but it is actually very popular in corporate world. Uh Manola remedy tests models for deep thinking. That's an interesting video. He actually compares uh big models and smaller models and he found that GLM 5.2 was very good. Uh and the Qwen 360 27B, very small model, was a close second. And he found that GPT 5.5 performed poorly because it was too agreeable. So, for simple task you don't see it, but for deep thinking uh you see that when the model becomes too agreeable, it actually hurts. >> [snorts] >> Okay, CoreWeave Area AI research agent. So, this is actually to help doing scientific research. Area can form hypothesis, trigger new experiments, evaluate results, update dashboards, and so on. It can process large volumes of experimental data and uh makes recommendations for the next steps in your research. AI research assistants. So, there are three listed here: SciSpace, Consensus, and Elicit. Uh they can do uh process a lot of literature, uh write literature reviews backed by verified clickable sources, which is very, very important. Okay, uh Qwabl 360 27B by Mia AI So, it's a fine-tuned version of Quinn 3627B for assistant style and reasoning heavy use cases. Focuses on step-by-step responses. Dario needs to be stopped. So, this is a very very interesting video. So, Dario Amodei, the CEO of Anthropic, uh so, they do the same tactics. They scare people that models can be harmful. So, Open AI is not good. Not Open AI, the company, but the idea that AI has to be open source. And he tries to convince that we need regulation by the government. And uh so, the author of this video provides multiple reasons why this is actually not true. This is bad. And that the only salvation, the only defense, is actually to keep things open source. Uh okay. X, formerly Twitter, released the MCP server. So, here are the links. There's a GitHub and uh some other links. So, you can use this MCP server uh to answer questions about uh Twitter you using Twitter data. But, if you actually want to use uh X data, Twitter data, you have to pay. So, the software is free, but access for data is not free. Okay, become more effective with Claude code. Uh so, this is a very good video which walks you like how to become a serious, efficient user um of uh Claude code. I will not go through this. Here's the link. Uh here is some advice. Create a brand context folder uh for your business, uh voice profile, define score voice rules, visual identity, positioning, use cloud code to point to these files, and so on. Uh next, Pulumi versus Terraform. Uh when you scale your AI application for multiple servers, you may want to use uh some sort of uh infrastructure as a code approach. The big gorilla in this area is something called Terraform by HashiCorp Corporation, which is actually uh here in America, in California. And the HashiCorp, I I think they're part of IBM now. Uh they created their own language, which describes your infrastructure. So, here's example of uh their language. You see, it's a declarative language, which has definitions. Uh so, here it describes the bucket and how to access it. And uh well, you have to learn this language. It's HCL. Uh kind of similar to YAML, a little bit. But, uh there is another approach, which is Pulumi. It's also describes infrastructure, but it does it in regular programming language. Now, for example, in Python or Go or TypeScript. So, here's an example of the same definition, but in Python. And as I'm a Python programmer, you know, for me, it's actually easier to read the the Python. And then you just uh execute this Python code, and it creates your infrastructure. Uh think about how you create infrastructure, let's say on Amazon Cloud. Uh you are on a terminal, you log in into Amazon, there is a command AWS, so you log in. And then using this command, you can create servers, create buckets, create file like you you can do what I whatever you want. And you can combine it into a shell script or in a Python script, but what Pulumi adds to it, it has templates how to do all the most common operations. So, for you, uh creating those objects becomes very, very easy. There are templates, there are examples, and you can ask uh your agent, let's say Cloud Code, "Just create me a Pulumi file to do this, this, and this." And you will get it in just few minutes. Okay? So, you can use your agent Cloud Code to create infra code. So, here is an example. Create a minimal Pulumi Python project for AWS. I didn't even specify what exactly I want. And uh Oh, no, no, no, here I describe it. Okay, not I. It's actually the agent created for me. And create the output. Do not add policies, whatever. And Cloud Code uses the prompt above to generate files such main.py, Pulumi.yaml, and often requirements.txt. And then you just run Pulumi up in the same directory, and that's it. It creates infrastructure. Very easy to use. Okay, jobs. Uh well, Microsoft just laid off about 5,000 people. Uh but other than that, uh nothing much happening. And this is me as usual, and thank you.

Links und Tools aus diesem Beitrag

46 weitere anzeigen

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Video

    Video:Lev Selector

    Wöchentliche KI-Updates: Modell-Releases, Prompt-Tricks und Agenten-Architekturen

    Lev Selector gibt in seinem wöchentlichen Rückblick vom 10. Juli 2026 eine Übersicht über aktuelle KI-Entwicklungen. Die Themen reichen von neuen Modellversionen (OpenAI GPT-5.6, xAI Grok, Anthropic Fable 5) über kostensparende Coding-Workflows und Prompt-Techniken für CLAUDE.md bis hin zu Fortschritten bei Spekulativem Decoding und Open-Source-Inferenz.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Neue Modelle, Agenten-Workflows und Infrastruktur

    Lev Selector fasst die wichtigsten KI-Entwicklungen der Woche vom 24. Juli 2026 zusammen. Zu den Schwerpunkten gehören aktuelle Modelle von Anthropic, Alibaba und Poolside, Kostenoptimierung bei API-Aufrufen, neue Agenten-Funktionen wie Claudes Skill-Recording, Google-Sucherweiterungen sowie der Umstieg europäischer Behörden von Windows auf Linux.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update vom 19. Juni 2026: Modellpolitik, Agenten-Engineering und Fusion

    In seinem wöchentlichen KI-Überblick fasst Lev Selector aktuelle Entwicklungen der KI-Branche zusammen. Themen sind der Rückzug von Claude Fable 5 nach Sicherheitsbedenken der US-Regierung, neue Open-Weight- und Open-Source-Modelle wie Kimi K 2.7 und GLM 5.2 sowie der Trend zu Ensembles via OpenRouter Fusion. Zudem geht es um Milliarden-Deals von SpaceX und DeepSeek sowie Architekturparadigmen wie Loop Engineering und strukturierte Speichersysteme jenseits von klassischem RAG.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Claude Opus 4.8, Agenten-Orchestrierung und steigende Enterprise-Tokenkosten

    In seinem wöchentlichen Überblick analysiert Lev Selector aktuelle Entwicklungen im KI-Bereich Ende Mai 2026. Im Mittelpunkt stehen das Release von Claude Opus 4.8 mit dynamischen Sub-Agenten-Workflows, Cursor Composer 2.5, die wachsende finanzielle Belastung von Unternehmen durch massiven Agenten-Tokenverbrauch sowie architektonische Lösungen wie MCP-Tunnel und Multi-Modell-Routing.

    KI & AI· News

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: Neue Modelle, Agent-Harnesses und Entwickler-Tools (August 2026)

    Lev Selector fasst die wichtigsten KI-Entwicklungen von Mitte August 2026 zusammen. Schwerpunkte sind neue Modell-Releases (u. a. Muse Glimmer, Grok 4.6, DeepSeek V4 Pro, NeMo Tron 3.5 Lightning), modulare Agent-Architekturen wie Prime Agent und DeepSeek Harness sowie Werkzeuge zur Token- und Kostenoptimierung beim Agent-Coding.

    KI & AI· News

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.