Vollständiges Transkript anzeigen (4.561 Wörter)
Artificial intelligence updates every Friday at 2:00 p.m. Eastern today, Friday, July 10th. >> [gasps] >> Uh after our July 4th holidays, 250 years of United States. And as you see, a lot of updates. And uh the epigraph for today is first we taught machines to answer. Now we teach them to act. So this is going from chat to agent. And actually the third step is not just to act, but take over our businesses. But we'll talk about it. So uh LM Arena. And this is from July 1st. But I just found like 5 minutes before the presentation that the leaderboard was updated. So let me uh just go there. At least for coding. So so this is uh the leaderboard, the original one. And what do we see? We see Claude, Claude, Claude. Uh so Claude is still on the top. Me use Spark. So this is from Meta, and they just updated it, made it more for coding and agentic work. And uh you see preliminary, that means that's maybe not enough data accumulated. Uh Claude, Gwen 3.7. Again, Me use Spark, the previous version. Uh Claude GLM. Uh Gemini Claude. So it's it's basically the same. Okay. So let me uh return to the uh hold on, presentation. And I go next. Um so this week we have a lot of new uh models released. I say new, it's not they're new versions. Like Grok from X A I space X A I. Uh so this is model which they claim is on the level of Claude, but it is cheaper. So you see it's a $2 in $6 out. This is per million tokens. Whereas the Claude Opus is 525 and Fable is like 1050. So this is of course much cheaper. It's a big model, 1.5 trillion parameters. And they acquired Cursor as you know, and Cursor has a lot of data for coding. So they use this data to train the model. Uh supposedly, I mean they're only first reviews, but it looks very very good. Um OpenAI GPT-5.6 models were released and people are extremely excited about them. So the Soul model, the most powerful, is $5 in $30 out. And it is their flagship, and review reviews are ex- excellent so so far. Uh so it's a Soul, Terra, and Luna. Soul, Terra, and Luna. Uh Gemini, uh during the IO event in May, they released Flash model, which is very good and affordable. And they promised to release 3.5 Pro, but they were delaying the release from June. Now the next day they promise is July 17th, so it's not out yet. It's in a preliminary preview. New Spark from Meta. So it was released yesterday, and it was already very good, and now with improved reasoning, coding, and video captioning, it it should be even better. Now, uh Fable 5 and Mythos 5 models were released. Well, Mythos 5 is not for the public, it's only for close partners of the government, but Fable 5 was released. And it is very good. I was using it for multiple things. It's amazingly good model. I have a Max subscription Claude Anthropic Claude Max subscription. And half of my allowance I can now use for Fable and so far I didn't hit the wall and yeah, it's amazing model. Very very helpful. So they currently have some promo, but it ends July 12th. So right now the promo is in effect and I have subscription and I can use Fable as part of this subscription. But after that starting July 13th, it will no longer be part of the subscription and if you want to use it, you have to pay for credits. Basically pay API costs, token costs. Okay, next. Instruct Claude to modify its own Claude MD. This is genius. It's simple trick, but very very effective. So what you do, you go into your Claude MD file. So we're talking about Claude code which has Claude MD file. If you don't have it, create one and every time you ask Claude to do something, it uploads this file. So it's always there, this file. And you add this instruction. When I correct you or catch you making a mistake before continuing, add the lesson as a one-line rule under the lesson sections in this Claude MD, so it never happens again. So you just add this one phrase in the Claude MD file. And then every time you correct it, it will remember it and it will become better at understanding what you need and how you want it. Like people use memory, people use like wikis, people use I don't know vector databases or graph databases or very complicated stuff. This is a very very simple like a one minute to implement and very very effective rule. So, this instructs Claude to update its own markdown file and when user corrects Claude, it automatically appends that mistake and a one-line rule under dedicated lessons section. So, you can create a lesson section in the Claude.md file. Anyway, Boris Chorney tweeted made a post on well, now it's X, not Twitter. But basically, he's saying that the previously we had roles like engineer and designer, but now the roles are more like prototypers, builders, sweepers who refine clean the systems, then growers which scale the systems, and then maintainers which ensure So, this is like five steps. And this is true and of course Boris Chorney is a genius. He is the original creator of Claude code, but I would say that this only covers implementation. But before starting prototyping, and it says generate many ideas. We need somebody like Steve Jobs, somebody with a vision, somebody who can define direction. And so, I write it here. So, we need somebody who knows where to hit, so to speak. So, remember this old story, a factory machine breaks down, consultant taps one spot with a hammer, machine starts, his invoice says $1 for hitting and $9,999 for knowing where to hit. So, we need these people and they're not like listed in this list. Another thing I want to mention, I really love Dan Sullivan. He has a company Strategic Coach who has more than 100 people, and he explains how business owners do everything themselves. So, they behave like self-milking cow, which is not healthy, and he teaches people to delegate uh to use the teams to achieve results. Um anyway, so this is uh real related to this. We need a cow. We need Steve Jobs. And then they can uh provide the milk or provide the ideas, provide the direction, and then all these other five uh types can work. Uh next, Open AI optimizations cut inference costs uh in half. So, you can serve ChatGPT for logged-out users with only a few hundred So, it's it's it's not for you. It's what they managed with the existing infrastructure and with existing mix of different clients. Uh but it's just interesting how being clever on using the resources, you can suddenly cut the host cost that profoundly. Now, DeepSeek-DiSparse speculative decoding. This is huge. Accelerates AI generation speed by 80%. Increases total output by 700% without losing quality. Uh so, they're using speculative decoding in a certain way using parallel draft enhanced with lightweight Markov head, eliminate errors, introduces confidence head dynamically, and so on. So, here is some technology. There is a video which explains it. Uh but what's interesting that this speculative decoding can be applied not only to DeepSeek but to other models. And this video, for example, describing using GLM 5.2 with DeepSeek-DiSparse, right? To for acceleration. And so, the result is you have a very powerful model which basically on the level of Claude or maybe beats Claude on some benchmarks. It's 85% faster because of DeepSeek-DiSparse. It's open source, and it's very affordable. So, that's what's happening. Okay. Oh, this is uh about my channel. So, I have 7,000 subscribers and 295 videos. Name of the channel Left Selector. I provide links for GitHub and Google Drive to download the slides, so you can download and explore all the links yourself. And I usually ask the question and please stop the video and answer the question in the comments below. Next, yeah, this this is a very interesting video. So, what the guy did, he tested three models, DeepSeek V4 Pro, GLM 52, and Fable 5, right? So, all very famous models and the task was to create a game, the Flappy Bird game. And they all, of course, created it because it's a simple assignment, but look at the total cost. This is Fable, which is the most expensive, about 42 cents. GLM 4.2, it's like about 5 cents. And DeepSeek is I don't know, it's like 1/10 of a cent. >> [laughter] >> So, it's it's not just per talking, but apparently DeepSeek also used less tokens to achieve the result. And here, this is actually a video in Russian, but what they were testing Claude Sonnet 5 against DeepSeek V4 Pro. And the total cost of creating website, so this was a creating website for a coffee shop. So, you go there and you select which coffee, which size, which like whatever you you have like a coffee drink configurator, right? Animated website. So, Sonnet spent $11 in tokens, DeepSeek only 8 cents. Again, Chinese models, that's why DeepSeek is almost everybody's favorite nowadays. We have our own agent and we using Claude and we using Deep Seek so we can switch. Deep Seek built a cleaner, more complete layout with a sticky menu, though its configurator had a minor pricing bugs. But again, you can fix the bugs very cheaply. Deep Seek will still be ahead. Organization-wide agent, Andrej Karpathy. So there was a presentation and tweet. And uh uh the idea is that first AI was a web chatbot. So you ask question, you get the answer, and the answer was mostly wrong, but whatever. Then we have desktop apps and uh now level three is persistent organization-wide AI entities. So desktop app, you can think of it as an agent which sees your files and can do what a human can do on your computer. >> [snorts] >> And uh persistent organization-wide, well, it started with Anthropic tag feature in Slack. So you can say at Claude and communicate with it. So Claude now knows everything that happens in uh in Slack in your company. So it knows your business and it participates after some time it becomes the most knowledgeable employee of your company and it can manage your whole company. AI [snorts] can now act as an active multiplayer employee deeply integrated across a company's entire system, tools, and context. Claude tag is currently only for teams and enterprise accounts in Claude, but in the meantime people already created similar technologies in open source on GitHub. Uh so a small business can build a similar functionality using custom Telegram bots connected via AWS and tools like Composer. So, this is about Composer and it connects agents to hundreds of apps like GitHub, Slack, Salesforce, Notion, Jira, Gmail, and so on. Handles authentication, sandbox execution, logging, dashboards, and so on. Makes it useful in turning agents prototypes into production workflows. Amazing tools, right? Uh Tencent HY3 open mixture of expert LLM reasoning coding agentic workflows 295 billion parameters 25 21 billion active per token. So, it's a mixture of experts. Context left 256 suitable for large code bases extensive documents fast versus deep configuration. Uh so, you can configure it to work faster or deeper. Strong on benchmarks cost-efficient for production use. Uh interesting. Uh next model routing helps to save 60 to 90% of costs. Everybody is doing it. Uh so, all the agents like open-source agents, commercial agents, they now offer you an option uh which model to use for planning when you need thinking and which uh model you use for execution, for example, for writing code. So, uh usually people use uh like a great model, let's say like Fable, to uh think about architecture and find the perfect solution and create a execution task task list how to build it. Not build it, but just make a plan. And then for actual execution, you use a cheaper model. And it may be again Anthropic Claude, I don't know, Sonnet, for example, or it may be GPT something or some other model. Uh doesn't matter, but because producing code, writing code actually provides a lot of output and output tokens are much more expensive than input tokens. So, by by doing that, by using cheaper model to write uh you can cut cost tremendously. Everybody is doing it now. Uh China, 140 humanoid robot companies, more than 300 models of robots. A major players, Unitree, Agility bought and the Ubitech. And they actually cover 80-85% of global installations of robots. So, China is definitely leader in robotics. Next, computational archaeology. Look at this scroll, right? With text. It is completely carbonized. It's like piece of carbon, right? And what scientist were able to do to decipher those scrolls. So, we have this 2,000-year-old carbonized Vesuvius scroll using high-resolution X-ray micro CT scan virtual unwrapping algorithms. So, they unwrap it as they look like a pieces of paper with some dots, whatever, and they decipher it using deep learning models trained to detect faint ink signals and convert it into text. And yeah, this is amazing work from single words to fully decoded scrolls. Uh Joe Rogan Experience Aravind Srinivas. Aravind Srinivas, this is Perplexity. I love Perplexity and he's amazingly brave, energetic, original thinking person. And this interview is long, it's about 2 hours, but I highly recommend to listen. They cover so many topics and he's like a living genius. Highly recommend you you interview. So, here's the link. Um As Syntax, he was silent for a while, but now he's regularly And what what he is doing, he is exploring different models. What you can run locally, what level of quantizations, and when you quantize the model, you can quantize different parts of the model differently. And so, his current favorite model to run locally is DeepSeek. DeepSeek V4 Flash. Very fast, extremely cheap via providers like OpenRouter, but he also runs it locally, and yeah, it is amazing. What he says is that for most of the task, well, at least half of the coding tasks, you don't need a big model at all. You can use a local model like 132B or 27B, and it does excellent job. So, you use a bigger model to think and create a plan, and then you use basically free model to do the execution, and this cuts your costs dramatically. Okay, Open Knowledge Format. This is coming from Google. This becomes very very trendy. I spoke about it last time. So, it is short format similar to skills, like Anthropic skills, but more for like knowledge records. And those files can reference each other, and this way you get interlinked linked system of files, and this is plain text. So, AI agent can easily read this files. It can, of course, find files and grep through them, kind of like MCP. And it works, and suddenly you don't need a vector database. It it becomes so simple, so light, so fast. So, OKF, which is open knowledge format, is a standardized file format consisting of structured markdown documents. Uh yeah, this is the way to go. So, it is much better than rug. Easy to implement, fast, understands the interconnections between different documents, the structure of knowledge. Um amazing how things become simpler. I remember uh rug several years ago, it was all vector databases, and now people do OKF with interlinks like a wiki, and that's it. That's all you need to have rug, retrieval augmented generation. Claude in Chrome is Anthropic's official Chrome extension. So, Anthropic released extension. Well, it existed before. This is, I guess, new iteration. It lets Claude see the page you open, understand its structure, and take actions like clicking buttons, following links, and so on. Combined with Claude code, you can build or modify a web app in the terminal, and then have Claude in Chrome open and exercise running apps to test flows, and so on. Functionally, it turns Claude from a pure chat assistant into a lightweight browser automation. Okay, so Chrome browser, and you have Claude extension. Uh use Python to make Karpathy style wiki. So, I just spoke about a file format, but this is a science and towards data science. It's a very good article. It shows how to do everything in Python. Which is close to my heart, because I love Python. Okay, how to talk to Claude? So, when talking to Claude to get correct answers, make sure to provide context. Because if you just ask the question, Claude may be confused. It doesn't know exactly what you're asking, and what your situation is, what your context is. So, you have to provide context, you have to tell it what is actually required and the role. Uh so, you ask Claude to behave uh like from which perspective it needs to think about it. Okay. KV, key-value cache, and paged attention. So, these are open-source mechanisms from vllm inference engine used to speed up LLMs. Yes, and vllm, by the way, is a very, very good technology and it allows to parallelize execution, so things run faster. During LLM inference, the prefill phase process the prompt, while the decode phase generates tokens sequentially. Uh okay, KV caching prevents uh by saving partial matrices. However, standard KV caching creates fragmented memory, so they used paged attention to solve this problem, cutting memory waste up to 80%. By the way, Syntax in his video, he also talks he because he was comparing different ways of running the models and he found that using vllm inference is the best. Okay, this video concludes with optimization tips including tuning memory utilization, enabling prefix caching, and using chunks prefill. Okay. Next. Uh Boston Dynamics fifth generation of Atlas. Uh so, you you know that they were acquired by Hyundai and Hyundai has a lot of experience in manufacturing, well, cars, but you see the robot now, they simplified the construction, they reduced the number of parts, they made it more technologically, so they are preparing to mass produce those robots and plans up to 30,000 units annually. Uh physical agility remains a core strength, but combining with advanced control systems, high-level decision-making, and so on. And this is another Chinese company, UBTECH. So, they create robots which looks absolutely like humans. >> [snorts] >> Uh full-sized humanoid robots for mass production starting at about $18,000. Uh lifelike silicone skin, motion joints, emotionally aware AI, and so on. Very interesting. But again, remember 140 manufacturers of robots and more than 300 types of models in China. It's huge industry. Okay. Oh, Chris Lattner, this is his picture. Uh he was famous for creating LLVMs, then Swift language for Apple. Uh and then uh his own company, Mojo language, which is kind of similar to Swift but better and for AI. And uh his company uh was acquired by Qualcomm. And Qualcomm is now using this Mojo language. And Mojo combines uh Python-like syntax, very simple syntax, and high-performance execution. So, you don't need CUDA if you have Mojo. And Mojo can work on different platforms like Nvidia, AMD, Apple chips, everything. So, this is very very exciting technology and Qualcomm takes advantage of it. So, by integrating Mojo with upcoming AI data center cards, Qualcomm provides cross-compatible software ecosystem capable of threatening Nvidia market uh chokehold. So, yeah, interesting. Mobile Open Claw now works with Android and uh iOS. Yes, Open Claw, you know, the famous um Open Agent. So, now it has ability to work on different kinds of phones, Android and iOS. Can talk to you, access photos, contacts, calendar, notification, device state. Okay, three AI world model paradigm. So, world model, you've seen those models where it's like a game where you see, I don't know, surroundings, people, houses. So, you have Meta's JEPA, which is Yann LeCun, World Lab taxonomy, and Einstein world models concepts. So, three different approaches. JEPA focuses on learning world within abstract latent space. World Labs structures world models into render, simulators, and planners. And Einstein world models framework from So, this is in Abu Dhabi in United Arab Emirates. So, it's Arabic world. And this is University of Artificial Intelligence. So, as modular tools, LLMs remain the primary reasoner. A quote on visual world models to perform thought experiments. So, the speaker argues that unified embedding architectures like JEPA may fail, predicting that future AI will favor modular tool-based like the Einstein world models. Okay. Now, this is a very interesting work. They explore how the JEPA actually works, and they compare JEPA to LoRA, and say that it's basically LoRA in disguise. So, LoRA, it's low-rank adaptation. That's how people do fine-tuning of models. Drawing a parallel low-rank adaptation in language models, the host introduces novel concept creating JEPA adapters. This specialized add-ons could efficiently fine-tune foundational world models for specific domains. Yes, very interesting idea. So, we have Laura for text models and here he proposes Jeppa for world models. Okay, um this is interesting article. It basically says that talented people using AI become more efficient and bad performance become even worse. So, that AI just amplifies what people have. Skilled engineers used to dramatically increase productivity, unskilled generate fragile code, technical debt, and costly failures. So, ultimately AI rewards strong fundamentals and punishes the absence, widening the gap between competent professionals and those relying blindly on automation. Okay, this is very interesting work. So, you see the person and this apparatus measures magnetic field on the surface of his head. Magnetic around the head using magnetoencephalography scanner. Uses AI to decode brain activity into text while the patient is typing. 61% average world accuracy peaking to 78 for top participant. But, it requires a giant laboratory equipment. So, as far as I remember those things are superconductors and you cannot have any metal pieces around it. So, you would put the lab somewhere in the forest far away from electric lines and like whatever. So, it's a very very specialized situation. But, just think about it. It can uh without any electrodes, without any invasion like from the surface measure well, not your thoughts, but predict the words you you you are typing. Holds massive long-term potential, definitely. Okay, uh 10 core components of an AI agent harness. Everybody now create their own agents and basically creating their own harnesses. So, instructions, context delivery, like reading files and so on, context management, filter compressed transcripts, reduce noise, tool interfaces, like schemas, MCP, execution environments, needs sandbox, container, durable state, persistence, memory, database. Orchestration, retry step-by-step, human approval, sub-agents delegation, skills and procedures, reusable playbooks, and verification and observability. Well, it seems obvious, but like I I look at this list and I realized we created our own own agent and we basically have all these pieces. They are required. You have to think about how to implement them. Cloud design was released the update, and it's much better. It adds a proper design system imports, better canvas editor, and tighter Cloud code integration. You can now use real components from a GitHub and design files, edit layouts, and so on. The goal is to keep work on brand while reducing token usage. So, for graphical designers, for brand designers, very good tool. Anthropic's Cloud certified architect exam. So, Anthropic now takes the path of Microsoft creating certifications. So, there are courses now. If you look, you will find many people offering courses, and you pay them to take the prep course, and then you can take certification exam if you want. Cloud is presented as a stateless model behind a single messages API with four key layers, raw API, agent SDK, Cloud code, and MCP. Okay, and the guide then walks them through each domain. Agentic loops, multi-agent orchestrations, cloud code configuration, wire cloud MD, prompt engineering with strict clickable checkable rules, robust [snorts] tool MCPU design, clear description structured errors, context reliability, and so on. So, yes, I think maybe it's actually useful to take the course like that. Uh well, I know this stuff practically simply because we were building things, but if you want to get into it, maybe it it's good to take a prep course because it will cover all all the details. Okay, Alex Hormozi, so he suggests three folder organization for AI agent. Business context, so this is MD file which describe your company identity, target audience, brand, and so on. Data, raw training materials like past newsletters, so this is about marketing for your business. And prompts, highly detailed reusable prompt templates. So, you create those three folders and maintain when you're working with your agent. OpenAI kills Atlas browser. So, remember every major provider of AI creating their own browser, and OpenAI browser was called Atlas, but now they decided next month they will shut it down in favor of browser extension and other technologies. So, shutting down Atlas after launch and redistributing into ChatGPT Chrome extension. Yes. Okay. Uh Georgi Gerganov, Bulgaria. So, he is the famous creator of Llama CPP. Llama CPP is GitHub repository, it's an engine to run models to do inference. We used used it also to fine-tune the models, but it was like 2 years ago. Now they removed this utility from this project. So they just concentrate on inference, on running the models. Accelerating llama.cpp using Claude Fable. Yeah, this is very interesting work. So the author, Kalakus, he used Claude Fable to optimize llama.cpp code to make it run like almost 65% faster. So it's very, very interesting. llama.cpp is used in so many places and like speeding up by more than twice is huge like worldwide. Uh very interesting. Uh open-source wrappers for Claude code. So this video uh introduces several things you can use with Claude code like for example Claude video, notebook LMCP, Graphify, Obsidian skills, Impeccable, which is open-source design tool, and Ponytail. Uh this is for token efficiency, so you can use two times less tokens. And this video talking about evolution of rag uh from simple semantic vector searches to uh modern rags. Well, he he talks about uh pre-retrieval and then post-retrieval processes. He talks about graph rags like structured knowledge. Uh which people used the graph databases, but it's really, really heavy and as I described, it's there's much easier way to just use MD files interlinked. Uh Agentic rag handles queries using autonomous reasoning loops. So you uh go and search until you find or you tell will I cannot find the answer. So, you using agentic loop. Okay, this is about jobs. We have a lot of layoffs in Microsoft. You see on July 6th almost 5,000 people and this is me as usual and thank you.