Vollständiges Transkript anzeigen (3.808 Wörter)
Artificial intelligence updates every week, every Friday at 2:00 p.m. Eastern time. Today is Friday, June 5th. As usual, a lot of updates. And this week in AI, like you see epigraph, everything got smarter, cheaper, faster, and somehow more expensive. Okay. So, let's look at our leaderboard. And uh this is from yesterday. Uh well, mostly looks the same. I want to point your attention to Zippo GLM 51, new version, which is pretty good. It's ahead of Google and and OpenAI, just behind Claude, right? Also, Muse Spark from Meta behaves well. And Qwen 3.7, the latest version, Max preview. Qwen also has smaller models and local models. Uh generally, Qwen is a good choice. Uh DeepSeek somehow doesn't behave well on this leaderboard, so it's all the way down. Anyway, uh Minimax M3, multimodal, text, image, and video inputs. So, it's a new Chinese model, version three, 1 million token context, long horizon, agentic workflows. And yeah, so this is a good model. Microsoft created a whole bunch of their own models. Like I think seven of them. Yeah, seven new in-house models. They're all called MAI, Microsoft AI. And this is Microsoft AI thinking. And yeah, it says that it beats Sonnet 4.6 good and software engineering benchmarks for multi-step agentic workflows. Okay, and it's available through Microsoft Foundry. Uh Uh, Nvidia, so there was a event. It was in Taiwan in Taipei, Taiwan. And there, as usual, there was a keynote by Jensen Huang, and he spoke about new hardware, new software, new models, everything. And Vera Rubin now is in full production. Uh, so this is architecture named after a famous astronomer, astrophysicist. And yeah, so so they have Rex, they have chips, they have all kind of stuff. Uh, Nematron 3 Ultra. Uh, they were very proud of it because it's open source, 550 billion parameters. And also Cosmos 3 open models. But if you look at Nematron behavior, like for example, on AI intelligence score, you will see that it's not on the top. And also on leaderboard, it's not on the top. But it is open source and uh that makes it useful. And Nvidia also created new chips for both laptops and desktops. And this is actually interesting. They're planning to release, like with Microsoft, uh, new laptops for Windows. And they have Nvidia chip uh, with 3 nanometer class, like a package. So it's a small chip. It's basically very similar to Apple silicon. Very comparable. Uh, but it has a GPU part, which can run big models at 4-bit uh, precision. So you know, models first were developed at 16-bit precision, then uh uh, how like compressed down to 8-bit. But this is actually allows you to run at 4-bit. So you can run pretty big model at a small chip laptop. You see AI performance up to one petaflop of a compute 128 GB of memory can run 120 billion parameters at a floating point four and 1 million contacts locally. So, this is this is really really interesting. Let's see what will happen in the fall. Uh now, this is on the bottom right comparison between Nvidia RTX Spark chip. This this is this new laptops and Apple Apple M5 Apple silicon chips. And they're very similar except for Nvidia has much more operations per second for AI. So, you see it's a big difference. Thousand compared to about like 60 at the top. So, so we'll see. Usually, my experience with Nvidia, they sell something very excited, but then it come out and it's not as good as it was promised, but we will see in the fall. Okay. Anthropic files for IPO. This is confidential. This is preliminary. Nobody knows the amount of this, but the latest Anthropic valuation was already close to trillion dollars. And uh Oh, this is quote of the day by Anthropic CEO Dario Amodei. Humanity is about to be handed almost unimaginable power and it's deeply unclear whether we possess the maturity to wield it. Okay. I made this graph just to see and show like here for example, Mercedes-Benz. You see or Ford or Honda. These are big companies, right? Toyota IBM IBM mainframes, whatever. IBM, I always thought it's a huge company. It's almost 300 billion, but compared with modern AI companies, Nvidia, Apple, Alphabet, Google, Tesla, Amazon Bedrock Bedrock is a library and infrastructure to run models including Claude of course, Anthropic open AI and and so on. So this is like 5 trillion compared in Mercedes to like 50 billion. So it's a 100 times more. Okay, and and and these are the the actual numbers for this graph. Okay, Anthropic new billing coming this June. Well, Anthropic again doing the same thing. They get us in love with their product and then they increase prices. So what will happen in June is the following. If you have subscription let's say pro or max, so $20 a month, $100 a month, $200 a month, you can still use it as subscription for let's say development. So if you using it using Claude Anthropic tools. So for example, if you use Claude code CLI in the terminal, it's considered to be interactive work and you don't have to pay extra for this. But if you're using it inside some other tool like a Visual Studio Code or Z editor or JetBrains or whatever, then these tools are using Claude via SDK and then it will be billed by token and that immediately becomes like 10 times more expensive. So you have to be very very careful after June 15th. What I recommend is turn off your API key and to see what works and try to work with subscription. And another approach is maybe to use other models for inference. So you can use Claude code Claude model subscription for doing development work, design work, but the tool itself should use different cheaper models. And here for example is a table showing some models, local or cloud, which are much cheaper. So for example, Claude Opus 4.8 pricing is $15 in $75 out. This is per million tokens in out, right? And this is a best model, but you see you can go much cheaper. So literally maybe 50 times cheaper, 70 times cheaper by just choosing So the quality will not be as good, but if you using it for smaller tasks, so you're not using it for planning, designing, but using it to just performing some regular task. Like for example, you don't need expensive model to parse your email or do some file operations and so on. So this is the way to maybe save maybe 90% if you go to local models, you can save a lot of money by doing that and develop your own application which works like harness, which uses expensive model for planning and cheap model for execution. Okay, so selecting affordable models. There are multiple websites which show and compare models like for example, artificialanalysis.ai and some others. I provide the link so you you can go. So this is from one of the images showing comparing. So vertical is quality and horizontal is cost. And here for example, Mimo V2.5, Deep Seek. You see pretty good quality and low cost. Minimax, this is Deep Seek V4 Flash. This is the recent one. Yeah, so you see how it goes. Uh, Qwen recent model 3.7 max and plus plus is cheaper and faster than max and quality is comparable. The These are in the cloud. Qwen max and plus they're in the cloud. They're not downloadable. Okay, uh, how to reduce agentic costs? Again, in the same area. So, this is very interesting wrapper called headroom. So, this is it has Chopra who created it. It's on GitHub. It's open source. And you start Claude like this. Headroom wrap Claude. Right? And what happens the headroom >> [clears throat] >> it's uh, between Claude and the cloud and it takes your context and it compresses it uh, reducing like by 60 you see to 95% reducing your token usage and it's very very easy to use. You just do pip install this headroom AI and then you just always start Claude through the headroom. Right? Uh, another this is a very interesting work. The idea of this work and this is on archive is that if you uh, created agent and agent is running some skills which are written as text files. Right? As so, it works but it's slow and every time the agent has to process those text files and use a lot of tokens. What if instead of that you will take this agent who is already trained configured with all the skills and you somehow trained a much smaller model a local model to do the same functionality. So, you literally use like 3 billion to 8 billion parameters Qwen model, very small model and you train it, and training process takes about an hour. And uh boom, you have something which is free, which is fast, and which is almost as uh precise as the original. And by proper architecture, maybe using several models in parallel, you can reduce hallucinations if you go for consensus, and get high quality and at the same time very low cost. So, th- this is a very good approach. And this is actually the website they're doing the same. It's called Open Universal Machine Intelligence. You see the right model with the biggest model stops uh renting generic intelligence at high cost, build and deploy the best small specialized model for your task, and own it fully. So, up to 50% accuracy versus frontier models, up to 90% lower cost, and you can train it uh maybe 1 to 2 hours uh from scratch to production quality. So, you see this is a common trend. People are doing this. Okay, this is for my uh channel, uh channel name Left Selector, which is my name. I have 290 videos and thousands of subscribers. You can always download the slides on GitHub or from Google Drive. I provide the links under the videos. And I usually ask question uh uh under the video, please stop the video and answer the question. Uh recently, it was all about how do you use agents, which agency use, what are your use cases. Um anyway, top five coding tools. Uh you've heard, of course, about Cursor, which was like what, 3 years, and they're now in uh relationship with Elon Musk's companies and valued at not at 30 billion, they're like 60 billion. Um GitHub Copilot, of course, big and Microsoft, well, I don't know its valuation, no standalone valuation, but it's a multi-multi-billion. Lovable is a big company from Europe, also like 6.6 billion valuation. Cloud Code, this is what we use, and again, it's part of multi-billion uh Anthropic business. Wind Surf, Aquarium, hundreds of thousands, low million of users, parent company valued at low single digit. Bolt, Vercel, Replit, Base 44, and others. But these are the large ones. >> [snorts] >> Okay, how to contain Claude across products? So, how to limit the blast radius when Claude makes mistakes? So, this is a paper from Anthropic, and they describe how they contain it, and how they use containers, sandboxing, and virtual machines. And Okay, and Deep SWE, this is a different thing. You know, so SWE bench, famous SWE, which people say it's not good. So, this is a better benchmark called Deep SWE. And you see, in contrast, SWE uses entirely original handcrafted challenges across 91 repositories, so five programs. It features concise behavior, and so on. On leaderboard, GPT-5.5 dominates, outperforming Claude Opus. Furthermore, proves highly efficient. And [snorts] I don't know, so it's a lot of conversation saying that the standard SWE bench is not good. It doesn't make sense. So, this this is alternative, and probably it's better. Okay, Anthropic SRT, sandbox runtime. So, if you go on Mac OS, there is actually a command utility uh which uh runs It's not like a doc docker. It's more limited, but it it allows you to put your application in a sandbox, so it will not destroy things which you don't want it to touch. Right? So, it's called sandbox exec. And it existed probably for about 25 years. In fact, now Apple has another mechanism which is uh well, they believe is better, but um Anthropic itself on GitHub there is a repository called Anthropic Experimental, and they provide a sandbox runtime. And this sandbox runtime, the way it works, it is written in TypeScript as Anthropic as code code is also written in in in TypeScript, right? So, it uh cannot control uh operating system level events, but it uses this sandbox exec command, which is on the level of operating system. And so, uh Anthropic TypeScript uh simply executes it with a certain profile configuration, and it works, and it protects your computer uh from AI. >> [laughter] >> Now, on Linux, there is similar utility which is called Bubble Wrap. Bubble Wrap is open source. You see it's on GitHub. And this diagram kind of shows on Mac OS and on Linux how you how you can protect your computer from AI. And this is a big problem, right? Because people complain, "Oh, AI went and deleted all my important files." Well, here how you can protect it. Uh next, uh Microsoft Scout. Uh this is an agent by Microsoft based on open AI or open claw, sorry. Open claw, remember it's an open source project which is now part of open AI. And so this is a Microsoft Scout is an always on autonomous AI agent first in new category called autopilots. Runs independently around the clock, integrates with Microsoft 365 Teams, Outlook and so on. So this is what Microsoft now offers. Now SpaceX IPO, there is a lot of conversations about it because they target 1.75 trillion in valuation but independent analysts says that the company is like overvalued and the actual value is not 1.75 trillion but only like 780 billion. So 1 trillion less. So we'll see what will happen. The IPO should happen middle of June. Next, JetBrains Melon 2 12B mixture of experts open source coding model. So JetBrains came up with their own model and it comes with three different variants which you can use in the IDE. Soonar, a very famous and fast growing company which generates music. I love it and they just raised 400 million at 5.4 billion valuation. Okay, next, five files to personalize your Claude code agent. So it's a Claude MD, Soul MD, Design MD, Voice MD and Audience MD. Who you are speaking with? I think it's self-explanatory. Next, Microsoft Project Solara. uh, chip-to-cloud platform. So, this is uh, for edge devices, and what's interesting, it's not using Windows. So, it's using a lightweight HOS called Microsoft Device Ecosystem Platform, which is built on Android open-source project, not Windows. So, they want something which is light, which is working on edge devices, which is doing AI. Uh, >> [snorts] >> yes. So, they're prioritizing lean flexible runtime workload Windows Okay. Uh, next, uh, agent compiler. So, this is again similar idea. You have an agent which runs some skills which are text, which is slow. And you want uh, somehow to optimize it, and today I already spoke about several ways of doing it. So, this is yet another way. It's called agent compiler. So, current agents essentially chatbots with tools, they re-reason from scratch every time they encounter a task. Agent compiler would transform agent learned experiences and reasoning patterns into optimized reusable execution paths, so agents stop wasting compute, rethinking things they already figured out. So, from interpret execute to compile execute. So, you're basically distilling the agent and making it like a smaller, leaner. Uh, the implication is significant for production AI systems. Rather than burning tokens, uh, redeliberating on solved problems, they'd run pre-optimized paths. Okay. Uh, Google DeepMind AlphaProof Nexus. This AI framework for autonomously solving open math problems. And uh, yeah, it already solved some problems. Remarkably, these breakthroughs cost only a few hundred dollars in computing power per problem. Uh this milestone proves AI can discover entirely new human knowledge promising dramatically accelerate future scientific So this is this is good. This is great. Uh on a this content layer MCP native context layer for AI agent converts emails, docs, chat into clean artifacts such as persona, voice, projects, and company knowledge and then shares via MCP. So you stay in control of your stuff. Cybersecurity recruiting is accelerating. Uh you know there were many scandals recently when AI was cracking finding holes in software. So people paranoid and hiring uh to do cybersecurity. Uh recruiters are urgently seeking leaders who can handle breaches, audit AI generated code, and secure infrastructure. Cybersecurity post job posting rising. Concerns intensified yeah because of anthropic my my methods and and so on. AI for consistent robotic work. Alibaba ships reasoning model that runs autonomously for 35 hours. Figure AI robots sort 88,000 packages okay in warehouse. Research papers and so on. It it it just to mention because usually I don't talk about it but that there's a lot of work uh around robots. Okay. Microsoft Majorana 2 quantum chip. So this is the chip. Thousand times more reliable than its predecessor, switched from aluminum to lead as a superconductor which helps shield sensitive qubits and making them more stable. Utilizes AI agents through the Microsoft Discovery platform to accelerate research. And they estimate commercial viable quantum computer in 2029, which is 3 years from now. Okay, Peter Young, five steps for building self-improving AI skills for Claude. Uh providing Claude with personal context, specific examples, followed by refining descriptions, implementing evaluation loop, and incorporating memory, which is very important. Very good video. Uh next, Anthropic when AI builds itself paper. Pretty substantial and important, and this is a uh Matthew Berman's video where it um uh is like discussed in detail. Uh he talks about recursive self-improvement. Well, he the the paper in AI development. AI systems increasingly being used to design and develop their own successors. A human role is transitioning from writing code to prompting and verifying, setting research direction. Bottleneck shifts from code writing to code verifying. Inside Anthropic, most of the code is written by AI already, and AI is significantly accelerating the speed in which tasks are completed. And [snorts] the three possible um outcomes, three possible futures. Uh this trend slows down, it's unlikely. Uh trend continues, but humans remain in the loop. And finally, uh people are completely removed from this development, and AI develops AI. So, Berman provides uh commentary on Anthropic's fear-based marketing strategy, noting that while the company advocates for slowing down the development of frontier to ensure safety, their current internal reliance on powerful increased models like the Mythus preview suggests they are highly motivated to maintain their competitive lead. So, the speed actually increases. Okay, Ray Kurzweil. This is Ray Kurzweil, um the famous predictor of the future. So, he still predicts AGI by 2029, which is 3 years from now. Uh citing continued exponential growth, he argues AI already outperforms humans in some domains like medical diagnostics and will reach vastly higher intelligence levels by 45. Uh and there's the discussion of some other interesting topics. The this is coming from a Peter Diamandis podcast on Moonshots. It's on YouTube. You can watch it there. Okay, uh context chain of thought. Uh context uh chain of thought enhancing context learning via high-quality reasoning synthesis. Okay, unlike standard in-context learning, which simply activates existing knowledge, context learning requires models to internalize entirely unfamiliar and contradictory rules provided exclusively with the prompt. Testing reveals that frontier models fail massively on these real-world tasks because teacher LLMs are lazy, generating flawed post-hoc rationalizations based on pre-training memory rather than new rules. So, context chain of thought solves this by hiding the final answers from the teacher models during data synthesis and optimizing stepwise reasoning. Implementing this approach yields a noticeable performance jump on difficult context benchmark. Uh well, chain of thought is not a new concept, but yeah, this is this is a good good work. Uh Huawei's semiconductor tau scaling. Well, there's a new chip design, a new metric. Uh it's a three-dimensional structure focusing on shrinking the physical transistor. Tau scaling optimize signal travel time, tau. And uh are optimize the corridors rather than shrinking the rooms. Uh logic folding, folding circuits elements vertically in 3D space. Huawei was using this approach over the last several years. Upcoming [snorts] Kirin chip expected in the fall will be the first commercial product to fully deploy the finalized logic folding architecture. Okay. Uh so, this is Huawei uh Uh next LLM interview questions. I don't remember where I picked it up, but I just put here some questions if you want to prepare for the interview about AI. Okay. Uh next. The AI layoff trap. So, this is a very interesting paper. And uh it argues that what firms do, they hire they sorry, fire people substituting them by AI, which decreases the costs. Which allows them to sell at cheaper prices and so on. And this is kind of a loop. It's automation arms race that displaces workers beyond what's collectively optimal. Ultimately harming both workers and firm owners by eroding consumer demand. And what the authors suggest uh that they need uh uh what's called Pigouvian automation tax uh rather than universal basic income or other. So, what is Pigouvian tax? The example is a carbon tax. A firm burning fossil fuels imposes climate costs on society, so the tax forces the firm to internalize those those costs and do something about this carbon emissions, making its private costs align with the true social costs, right? So, uh the paper argues that other commonly proposed remedies fall short. Like, for example, universal basic income redistributes after the damage is done, doesn't change firm incentives, right? Or upskilling, retraining, or capital gain taxes target profits broadly, not specific externally. So, yeah, maybe the peruvian tax is the way to go. So, the companies which substitute people by AI, they should pay some sort of peruvian tax. Okay, these are as usual updates about layoffs. Well, June just started. Oh, this is an interesting paper. A top economist says there is a zero evidence AI is killing jobs despite thousands of actual layoffs. So, it's an interesting analysis. Anyway, this is me as usual. And thank you.