Vollständiges Transkript anzeigen (2.991 Wörter)
The piece. >> [snorts] >> Artificial intelligence updates every Friday at 2:00 p.m. Eastern Time. Today is Friday, June 26th. Um, a lot of updates. Uh, so there are new models, uh, which was told by US government. And currently the top models are GLM 5.2 from China from ZAI. Uh, many people like it, although many people show that Opus 4.8 is still much better. GPT 5.5 Harness engineering becomes a new gold rush because with models becoming close to the qualities, the harness is the um, decisive thing. Anyway, uh, hold on a second. So this is leaderboard arena from yesterday. Doesn't change much. You still see that Claude is absolutely dominating, uh, the coding. Uh, Qwen 3.7, GLM 5152. These are Chinese models open source. Okay, so this is for coding, this is for regular text chat. So everything is about the same. Uh, Anthropic is still offline. I'm talking about Fable 5. And, uh, nobody knows when it will reappear. It may takes weeks or months. Unfortunately, uh, government kind of slows down. And [clears throat] they now did the same for OpenAI. OpenAI wanted to publicly release GPT 5.6. And government said, "No, no, security concerned. First, uh, release like to trusted partners one at a time, only approved by the government." Uh, recent Trump executive order that asks companies to give the federal government early access up to 30 days to the most capable models. So, government really want to take control and slow things down. Uh yeah. This is the famous story of Alibaba. The [clears throat] reason why I wanted to talk about it is because Alibaba was added uh to the blacklist. It was blacklisted in the USA. That means that uh uh government organizations cannot uh make direct contracts to Alibaba and next year they will not even be able to use any services of Alibaba. And I think the reason why is because Alibaba is now um controlled, maybe not completely, but uh maybe completely by Chinese government. Uh so, what's happening is in 2020 Alibaba founder Jack Ma gave a speech criticizing uh China's financial system. And um CPP uh Communist Party of China decided it's not good. And they basically they stopped him. They stopped the uh massive uh IPO he was planning in two places. They charged him with $2.8 billion antitrust fine. And they basically shut him down. Before he was open public figure and suddenly he disappeared. So, valuation of Alibaba dropped. His personal value dropped like in half from 60-something billion to about 30 billion. And he became almost invisible. Um so, this actually hurt China because uh investors uh they lost trust. They don't want to invest into Chinese business if they know that Chinese government may take over and shut it down. Um but anyway, this this is the situation. And you see US government has added Alibaba to the blacklist because it's now mostly controlled by uh CCP. Uh Z code and GLM 5.2 versus Anthropic. This is interesting video. Um many people now say that GLM 5.2 is a very good model. It's a model from Zippo AI from z.ai. Uh very high quality on the level of Anthropic with very close. But uh they were doing some tests using Z code, which is a Zippo analog to Claude code. And what they found that the model was not able to solve tasks, whereas when they were giving the same task to Claude code with Anthropic, uh it was solving it instantly. So, Anthropic is still the king. Although GLM 5.2 is amazing model and it can be used in and it's absolutely free. It's open source. It's free. Uh Sub Q sub quadratic. Um so, this is a new way of handling long context. Here, 12 million tokens context model. So, it's a Miami based startup called Sub quadratic. It has this proprietary model. It's hosted. It's available via API. It is very fast. And it's very cheap in comparison with uh frontier models. Uh so, for coding agents, analysis, very large documents, strong long context retrieval, significantly reduced attention flops versus attention transformers. So, sub q.ai it's a website. And it's available via API. Uh 1M preview and small model. Okay, this is sub-Q. If you see this abbreviation, sub-Q. Uh supersonic tsunami. Peter Diamandis wrote this article, and it has some really interesting stats. Like, for example, the token price was going down and literally like 1,000 times in just few years. People don't realize how fast the prices go down and how fast the capabilities grow. Um so, we went from 175 billion parameters in 2020, which was the peak performance, to now multi-trillion parameter models. Like, for example, Claude um Opus is estimated about five trillion parameters. Uh AI could push global GDP from today's tens of trillions towards hundreds of trillions or even quadrillions. Uh medicine will become much better. Uh AI compressing decades of drug discovery in diagnostics into like 5-10 years. So, this is very optimistic. Uh just showing how fast the progress going. Okay, this is plug for my channel, as usual, close to 7,000 subscribers and 300 videos. Uh name of the channel is my name, Lev Selector. I always provide links to slides on GitHub and Google Drive, and I always ask question under the video. Please stop the video and answer the question. Okay, a brain in the middle. Uh you know this picture? This gray is me when I was about 30 years old. This is Dim Boryakov, this is Sergey Nikitin, and together we created a small medical device. Uh so, this is kind of a prototype. This is I think these are uh whatever, they were big uh electronic devices cost from like $30,000 to maybe $100,000 to measure electrical signals from human body using needle electrodes or skin electrodes. And the way it worked, so you have some input uh from electrodes, you have some custom electronics so in inside a lot of transistors, and then we have some outputs on the screens or on paper. So, what we have done, we threw all these electronics out and instead using computer brain. So, we use software. So, we still need input uh amplifiers and filters, and we still uh well, the output was on the screen, right? But, the whole thing simplified because we used software in the middle. And uh now we can do the same, but using brain as an agent. So, large language model agent. So, uh if you can see the software classical application, you have some inputs, you have some functions, interfaces, whatever you created. Uh let's say software is a uh SAS system, and then it produced some outputs, actions, reports, decisions. Now, you can do the same thing, but you will throw out all the middle, and you just put agent. And agent uh you can customize it with skills, and uh it becomes very, very generic uh how you create applications, and very easy. Much easier than to program the custom software. Okay, Sakana AI Fugu ultra orchestration. So, Fugu systems um match or outperform Anthropic Claude Fable and Mythus on some benchmarks. Uh pricing is about the same uh as uh with uh Claude. Um and they also have pro and max subscription. Uh and the performance comes from coordination of multiple models. Uh behaves like single frontier class endpoint, but architecturally it's an orchestration layer rather than standalone model. Um [clears throat] so Sakana AI photo um ultra architecture Uh multimodal orchestration systems. So, not LLMs, but orchestration systems. Dynamically coordinate multiple AI agents through a simple single API rather than solving task in one monolithic model. Built on Sakana's Trinity conductor research, the system handles model selection, delegation, and synthesis internally. And uh yeah. So, now it's possible to create very capable model by combining uh multiple models together. And there are multiple solutions for that. This is one of recent ones. Okay. Reflection AI, uh famous startup. Uh so, now they made a deal with SpaceX to use their Colossus uh two data center. Uh paying 150 million per month. Please mute. Who is not muted? Ilya, I'll mute you. Okay. Uh Claude Corp fellowship. Anthropic recently unveiled uh a year-long fellowship. Uh places early career workers inside nonprofit organizations to help them incorporate AI. So, interesting. So, AI goes into nonprofits and Claude uh Anthropic pays for that. Introducing those fellowships. Uh Grok raises uh 650 million. So, Grok is a hardware. It's a chip. It's servers and 6 months after Nvidia license is the technology and poached the CEO. The company, the remaining of the company now raised money to pivot from chipmaker to inference-focused cloud. Okay, Base 10 raised 1.5 billion. What Base 10 is is a flexible infrastructure. So, they use other people's hardware, for example, Google Cloud using Nvidia GPUs and they allow you to create your own platform for deploying, optimizing, and so on. It's a privately held founded in 2019, San Francisco, California. About 60 employees. Interesting, the growth is very fast. Okay, next, Harrison Kinsley. His YouTube name is sentdex. I really, really like his him his channel. He historically was making a lot of tutorials on Python and JavaScript web development and other stuff than machine learning. And in this video, he talks about his recent experience with AI. And this is his custom built based a computer at home. And you see this Nvidia and then another Nvidia and another Nvidia. So, he has four Nvidias and each one is like more than $10,000 and 96 GB. And he was running GLM 5.2 on this desktop. But he had to use the quantized like 4-bit version of it to be able to run it. So, he talks about different models and Yeah, they're testing different models how good they are. I actually recommend you to to listen. I I like how he talks his emotions, his stories. And yeah. Uh Hermes agent skill learn. So you can say learn and then tell what to learn. Uh you can provide URLs, you can provide directories with files, you can provide whatever sources you want. And it will learn and create a skill. So here for example, learn how I just deployed the staging server focus on rollback and verification. So what it will do, what Hermes agent will do, um it will look at your chat, it will look at the history, it will see what you were doing for deploying whatever and it will create a skill out of it. And it will call it for example deploy staging and then later when you want to deploy something, you can just use this skill. So makes it very convenient. Uh what people do now, they use multiple agents in different architectures and the most common architecture is when you have a central like main agent which is orchestrator or conductor like in a musical conductor like orchestra, right? And uh it gives assignments to multiple sub agents and then sub agents send back the results. So this is kind of like map reduce in Hadoop. Uh same thing and uh multi agent orchestration agent map reduce. So it's a well it's known by different names like LLM map reduce, LLM map reduce pattern, agentic map reduce, uh map reduce for agents, whatever. And uh yeah, just interesting how the ideas from originally from Google, how they were doing Google search and then Hadoop, Spark and now multi-agent systems, which may have hundreds or even maybe thousand agents. Uh Seedance 2.5 is Bytedance's latest AI video model. Uh it supports to 4K uh output around 30 seconds. Uh so, it's a latest and greatest visual model. A Rust-based rewrite of SQLite. I love SQLite. It's not AI, but >> [laughter] >> I I I just want to put it because it's very, very important. SQLite now is part of standard Python library, and the most recent version is you see this project uh is written in uh Rust. Although, I'm not sure if this is part of standard Python library though. It's maybe separate project. Okay. How Anthropic own team gets AI to stop lying to them. So, this is a video where he goes um how to make sure that uh the model is uh truthful. Be specific, create a single source of truth. Uh require receipts, get a second opinion, test before automating, and so on. So, it doesn't happen by itself. You have to do certain things to make it happen. How Anthropic employees use Claude skills. Yet another interesting video. Uh so, Claude skills are structured for users of scripts, assets, and docs. Organize skills into four types: utility for specific types, verification for quality checks, data enrichment for external information, and orchestration for chaining steps. Uh components, verification, gotchas, and triggers. Anyway, very good video. Highly recommend to watch. Uh Fable brain prompting. Fable brain prompting. Oh, this for a image. Um anyway, uh from official Anthropic developer documentation that governs Fable behavior and compiled it into single system prompt. Act, do not overplan, lead with the outcome, ground every claim, stop only at real boundaries, assess, do not act uninvited, give the reason, provide the context, match effort to the task, and keep lessons, and check your work. So, these are things you can ask your model to do, and then its behavior will be very similar to what the the best models like Fable uh present. Uh the limits of prompting, the presenter emphasizes that this approach works for communication. It cannot create judgment. When given a flawed task, the Fable-trained model may still provide a confident but wrong answer. Okay. Harness engineering, a new gold rush. So, what is shown that by using different harness, you can uh with the same model up to six times better results by just doing using better harness. So, how you create better harness, and people creating their own harnesses. Like for example, sent text, uh which I uh was talking about before, he created his own uh harness uh for for his own tasks. And yeah, you can do a context management, treating memory as hint to be verified, connecting tools to validation checks. Uh so, we all need to start creating our own harness, and test it with different models. Uh build a custom agentic OS. Uh you probably have seen often recently agentic OS. Uh skills and automation, memory and state, uh interface, visual and distribution how you package and send it to users. So now people in business of creating this agentic OS and provide it to their clients. Scheduling for your digital employee. We have created our own agent and as part of its work we need to automate like schedule certain tasks. And for example, let's you want schedule something in your browser using your own account. So you want to use your browser which already logged in into some accounts, right? You cannot schedule it outside in like let's say crontab because then it will not be logged in. So you create your own script which is in your account using your open sessions, right? Fire agent skills not just shell commands. One demon for all users. We can create multiple users in the same. Okay. Hot reload on edit, on demand reload, random delay to prevent thundering many requests at the same time. Session persistence for recurring task. Let's say you want to run something every hour. Per minute deduplication prevents double firing, isolated per user login. So I've just listed some of the features which we experimentally bumped into and fixed. And Vitaly, thank you for doing all this work. And it took a lot of experimentation, but this is actually how the real system works and you see there are a lot of things to think about and to make happen. Okay, three layers of AI agent. So let's say you created an agent and now you want to give it to a client and then you want to give it to another client and then you want to give it to another client. And they wanted to have some customization and you realize that you don't want to maintain multiple forks of the same agent. So you want to have some layered structure. Like think about operating system. You have a kernel and then you have some applications, some tools and you flexibly package whatever you need to deliver for each particular client. So for the agent you may have a stable core runtime which reads manifest, loads plugins, manages MCP, lazy loading and so on. Then we have some shared plugins and then you have uh specific configuration for particular client, right? And so you have a at least three layers of structure. Then everybody may get the same core but then download plugins as needed and configure as needed. So this is actually what we're working on right now. Okay, Peter Diamandis and you know his channel about moonshots, they talk about recent developments. So here his vision as we get more and more sensors by 2030 tens of billions of connected devices with trillions of sensors. So creating near total visibility uh on the physical world and radical transparency. So people will know that they are constantly watched and probably they will behave differently, but this of course privacy and power balance concerns. Um GLM 5.2 success. Everybody talks about GLM 5.2 from Z AI. Uh Orion token price index tracks the cost of AI inference treating intelligence as commodity. A similar to oil. It's a interesting price index. Okay. Uh, well, not many layoffs in June as you see, much less than it was in May. And this is me as usual. Thank you.