Vollständiges Transkript anzeigen (3.470 Wörter)
[sighs and gasps] >> Artificial intelligence weekly updates every week Friday 2:00 p.m. Eastern Time. Today is August 7th. A lot of updates as usual. And first the leaderboards Well, nothing changes much as you see Claude dominates both coding and just regular chat. Claude is blue. Red is Gemini. Yellow is OpenAI and uh green is open source, which [clears throat] is mostly Chinese. Notice Muse Spark from Meta. So, we have Muse Spark, Muse Spark 1.1, and Muse Spark 1.2. This week they came up with a version for coding. Anyway, next This is a cost per intelligence. which I show every week and you see for the same task Claude Fable 5 is very expensive like $3, whereas DeepSeekV4-Flash is 3 cents. And GPT-5-6 Luna is also very inexpensive and good choice. But this is the link artificialanalysis.ai. There are many other charts there, many options. I highly recommend to look at it from time to time. It's actually not July 21, it's today. But anyway, >> [clears throat] >> xAI launches Grok build mode. You know that Grok has Grok build. It's like Claude code. But now they have it in the chat and you can build applications. So, it's basically text to app. Allowing users to create apps text prompts Drok me and put them with a URL so people can use this application right away. The system integrates coding agents, image generation, life web browsing and and so on. Multiple parallel helper agents. So this is a good thing. You can very quickly create a web application. Sakana AI nature-inspired intelligence. Sakana, it's a Tokyo-based company. It's amazing, very innovative. And they came up with many ideas like model merging, autonomous research agents, biologically inspired architectures, continuous third machines, and so on. Spike neural networks, this is like neural spikes. So it's not just connections, but these are spikes which are running in time like in real brain. Code-based memory MCP, fast and token efficient MCP server that indexes code bases in a persistent knowledge graph for AI coding agent. There is a trend now that people create for example Obsidian Wiki. So it's a MCP files which are all interconnected, right? And this is basically a simple graph. So different versions of graphs. So this is a code-based memory MCP parses source code with a tree sitter and stores structural relations functions, classes, calls, imports, and routes in SQLite. And agents can query graph instead of rereading files, which can cut token usage dramatically. So nothing new, just a good implementation and it is on GitHub and it's generic. Now, Demis Hassabis, famous creator of uh uh DeepMind and which part of Google. Now he is becoming a chief scientist of Alphabet, which is mother company. And yeah, that's I guess it's a big promotion. And yeah, he's also a Nobel Prize winner. Just to remind you amazing person. Thinking machines, they have the model Inkling which is close to billion parameters and now they creating Inkling small. Which is 276 billion parameters. So it's a four times smaller. Nearly the same performance. Context length 1 million and it's very inexpensive. So this may be a very good model to use. So this is just showing this is Inkling small and this is Inkling. It shows total number of parameters and the dark is the active parameters. So you see it's very very small can run on modest hardware. Google Gemini Spark AI personal AI agent. Operates on Google workspace. So it's Gmail, Docs, Drive, Calendar, Sheets. And it [clears throat] has Chrome auto browse. So it can pilot web browser to do research and so on. So this is a Google side of things and it is good. OpenAI Astra solves 10 math problems. Yeah, this is absolutely amazing. So OpenAI announced Astra and unreleased long horizon model family built specifically for parallel multi-agent execution on complex multi-hour technical assignments. And they did in fact solve long-standing mathematical problems. So it's not about doing it fast, it's just about doing it, achieving it. Solutions were generated for roughly $2,000 in API compute tokens with every mathematical proof formally verified using machine-checkable lean certificates. Very interesting. Um Muse code versus Claude code. So, I told you about Meta, which is Facebook, uh creating this Muse family of models, and now they created a Muse Spark and specifically Muse Spark 1.2 model, which is focused on long horizon software engineering. Right? And Muse code, you see it you install it like Claude code, and you you use it in your terminal. It's like Claude code, but specifically for Muse Spark model. So, it's a harness, and it has automatic fanout of parallel subagents, and so on. So, positioning versus Claude code. It's not as good as Claude code, but it's not expensive. And so, it's like a fast, cheap worker, which you can use. Now, about Claude code, Claude code now has a new feature, /doctor. So, this is a command, you run it, and it runs health checks on your installation, configuration, skills, MCP servers, and tell you what's wrong, where you can safely trim something or fix. You just /doctor, or you from Unix prompt can run Claude doctor. So, it is a built-in feature of Claude. It's not third-party. It's from Claude itself. Okay, this is about my channel. So, again, it's on YouTube. Name of the channel is my name, Lev Selector. More than 7,000 subscribers, 300 videos. I provide slides. Under each video, there are links. Please pause the video and answer pinned questions. Uh And one more thing, I started making uh uh with Elena Potapova, we started making short videos. So, now there are shorts on my channel, which you can check out. Um Claude watch skill. Very interesting thing. It actually can watch the video and summarize it. Okay? So, it explains a little bit how it works. So, it takes the frames and analyzes it. Uh so, you can use it. Okay. Buzzy, a cinematic AI canvas. Uh somebody is not muted. Please mute. Uh Agentic infinite canvas for end-to-end AI video production. Integrates more than 50 tools, more than 70 image video models, uh including the best like C dance, video clean, Runway, and so on. And offers director level controls. So, this is a pretty sophisticated system to create videos. And you see the link is buzzy.now. Uh next, Agent OS by Julian Goldie. So, what it is, Julian works with many different agents and he is into marketing, social media. And when new agent or new model appears, he wants to start using it. So, he created a system where you can add and add and add different things. They may be models, they may be agents, harnesses, whatever. So, he has a mission control dashboard where you can see all your models and all your agents. Uh he uses Obsidian type wiki with markdown files, which can be shared between those systems. It has a router to choose which model or which agent to use for which task. So you can look at your agents and the workflows and it it has its own loop engineering for continuous improvement. So once you're running some tasks at the end it analyzes what happened and tries to improve the skills. So some of the things he's using agents by themselves and they have similar mechanics inside them. Some I I I just models. But this is like a I don't know, combining different things together in one dashboard. And he actually says that anybody can do it for you can do it for yourself in like half an hour. You can for example, you can ask Claude code to create this local application for you. It will open in the browser and you will see all all your installed systems. Now Hermes agent and Buzz team workspace. Hermes agent what is the same thing twice? Why did I do that? Anyway, so Hermes was updated and now they have Herald release which is finding its voice. What it is you can talk to it now. It's a real time voice and it's a bidirectional. So you as it thinking you can continue talking to it. Voice notes now work across apps like WhatsApp, grounded citation, safe approval prompts. Open line idea. So continuous collaboration. You can continuously talk to Hermes now. Tailscale connects local AI agents to mobile devices. End-to-end encrypted mesh network without exposing ports. So you can talk to local devices. Liquid LFM 2526B small local agentic model and runs in a very small amount of memory so you can work on phones, laptops and on CPU only. Liquid [snorts] AI says it work out of the box with harnesses like Hermes agent open claw and pie. Models trained with multi-step agent workflows. So these are the links. Minimax uh H3 so this is video AI generates video from text, images and audio. So well it's a good thing. Next Neuralink wheelchair and brain computer interface. Uh so Elon Musk's company now uh market valuation surpassed $40 billion so for several years they were developing uh interfaces between brains and computers. Uh then they were doing this on monkeys then started on humans but humans were controlling uh let's say cursor on the screen. But now it's actually controlling a wheelchair. So a person sitting in the wheelchair can navigate this chair with his brain. Uh so person cannot move his limbs but he can move the chair. Uh very interesting. Qwen 38 marks open weight model uh uh it's available. It's Alibaba's best like the flagship model they have 2.4 trillion parameters for general use pricing $2 in, $6 out for million tokens so very affordable and you see uh it is blue here. It's yeah very good performance across multiple uh benchmarks. Uh Dreamina Cidence 2.5 AI video generation. So, uh Dreamina is ByteDance AI creative platform. Cidence is the video model, and CapCut is ByteDance editing application. So, ByteDance is a huge uh company in China. And this is um the frameworks to work on videos. Uh Nine Router, free AI router, and token saver. Connects locally to many different models uh in different providers, integrates tools like Cloud Code, Corsa, GitHub all together. Uh built-in optimization features to reduce token usage. So, you can run local and uh build complete applications without recurring costs. So, you see the screen shows many, many things it may have inside. AI system for endless content ideas. So, this is interesting video where uh the creator discusses how you can generate uh new ideas for your projects. Uh next, Tencent open-source AI memory system. Uh large context window degrade LLM performance over time causing AI agents to waste tokens. Tencent released open-source MIT license uh memory plugin described to cut token overhead while drastically boosting benchmark success rates. Compresses raw execution logs into interactive mermaid diagrams. So, this is really interesting. So, now uh the memory is not just text, but also mermaid diagrams. Inspired by cognitive psychology, it structures memory into episodic events and four semantic layers ranging from facts to persona defaults. Uh running locally via SQLite provides fast hybrid search, complete privacy, and zero API vendor dependencies. Uh GraphRAG, uh well, GraphRAG originally, this is from Microsoft, and this video and GitHub, their positioning is that graphs are better than vectors. And yes, like 2 and 1/2 years ago, we were playing with the vector databases for retrieval augmented generation for RAG. Now, we're just using uh wiki type structure, which is basically a graph, and we're getting better performance. Uh next, uh uh Julia McCoy. Um here, you see she's wearing this device, and so what are these? These are AI-enabled devices, so you don't have to be a slave for your desk, like sitting in front of your computer all day. You can walk around, and at the same time, you can talk to your AI, and you can direct it to do things, or you can talk to your team. So, her position is liberating from desk work. So, this is a new era, and this is how she lives. She really lives like that. She has a team of 15 people. She used to have 100 people. Now, it's down to 15. And yeah, she doesn't have to sit in front of the computer every day. Um Abacus AI self-improving auto bots demo. So, this is this video. Uh auto bots inside chat LLM self-evaluating agents designed to continuously grade their performance and autonomously optimize their workflow strategies over time. Uh example of usage, sales lead scoring, stock paper trading, YouTube thumbnail testing, and so on. Abacus is a very good company. They have a product which is called Deep Agent, and they uh well, they it's a commercial company, but they high quality and they run their agents on server. They are SOC 2 compliance for big enterprise. So, this is actually a very good product and very professional knowledgeable people behind this company. Microsoft Fara 1.5 browser agent model. So, you see Microsoft Fara. So, it's a family of browser automation models. They're small, 4 billion, 9 billion, 27 billion built on Qwen 3.5 with open weights under MIT uh license. And so, it operates observe, think, act loop enabling interactive decision-making across multiple step web tasks. So, this is working with web, supports long workflows, relatively long context for the small models. Uh includes safety checkpoints, best used in sandbox environments, smaller 4B model run locally efficiently. Okay. Impera Qwen 27B, which is uh based on Qwen 3.5 27B. So, it's open, local, multimodal, 1 million context, open weight, Apache uh vision, images, screenshots, charts, handwriting, native multi-talking predictions. There is also a smaller model, not 27B, this is 9 billion parameters. And uh Impera is independent research lab, probably in Germany as as far as I found. Okay, yet another CatCoder version 2.5 dev. Apache open weight, mixture of experts for coding from China, uh based on Qwen 3.6, and trained like with 127,000 supervised examples. And yeah. For the small model, it provides very good results on benchmarks. And, yeah. So, it's a good model. ZLUDA, running CUDA on AMD graphics. ZLUDA is an open-source software compatibility layer that allows unmodified Nvidia CUDA applications run on AMD uh graph- graphical cards on AMD GPUs. Translate instructions on the fly. Uh and so on. Pi on RTX 3060 12B. So, Pi is a lightweight coding agent hardness for local models. And, as you can see, it can run on a old and small GPU memory uh card. And, they discuss prefill speed and uh token [clears throat] per second generation. Prefill speed is the starting delay. So, time to first output token. And, uh for local run, it's actually very important characteristics. So, mixture of expert and grouped query attention can reduce VRAM and KV cache pressure. Many Llama CPP guys are outdated. A small focused tool set can make local agents feel much faster and cleaner. Uh local setups win on privacy and control, but still lose on hardest long horizon tasks. Yes. So, this is how you run when you don't have a lot of memory. Uh Rust quite system takeover. >> [clears throat] >> So, here, for example, for Discord, you see memory usage from 8 GB down to on- only 400 MB. Latency spikes from 300 ms to almost zero. CPU usage from 100% to 60%. Latency [clears throat] from 250 ms to 10 ms. So, this is just one example. Uh government American government uh uh they push uh their contractors and to start using memory safe languages instead of CC plus plus. At first, they were kind of mandatory, now they kind of relaxed shifting to risk-based agency discretion instead. Meanwhile, Rust adoption continues. Everybody is using it. Google cut Android memory safety bugs from 75% down to 20%, mostly because they removed CC plus plus, Linux, Android, major cloud vendors, and so on. DARPA, which is military, and other explore AI assisted C to Rust migration. Lots of examples. Rust truly is becoming very important language. So, if you are an engineer, you have to know Rust. You you just must. Gen Office, Gen Spark. So, there is a company in Palo Alto, California, and they created this workspace, which is very similar to Google Space. And it is AI enabled on top of that. And for example, they have Gen Office, which is like Microsoft Office. So, you see here sheets, docs, and so on. So, it can work with Microsoft Office documents, but it is Gen Office. It's open source. It's free. And it was actually initially created by a single engineer in about a week. So, this was some time ago. So, nowadays, I think any one of you can just talk to Claude Code Fable and ask it to generate something like Microsoft Word, and you have your own Microsoft Word for free. So, you probably don't have to buy Microsoft Office anymore, and you don't need all the features, but if you need to add some features, you can always just ask to add those features. So, this is a true revolution. >> [snorts] >> Um so, GenSpark is this workspace. It's all-in-one agent workspace. Uh but, GenOffice, so this is genoffice.ai. You can just go there, download it, and start using it. Uh very interesting project. Okay. Uh MobileNext, uh this is a development tool set for people who create mobile applications. So, uh imagine that you uh were doing development on your laptop, and now you want to test this mobile application. You can do it on a emulator on your laptop, or you can connect a live phone, and you can run uh on this um connected phone. So, what this mobile CLI does, it can effectively press on the screen of the phone, do the swipes, do whatever. So, you you can test your mobile applications, and uh you can describe using text. So, it's AI-enabled. You can tell it what kind of tests you want, and what to do, and it will do it. So, there is a mobile MCP, mobile right. It's like Playwright. Uh if you're familiar with the controlling browser on the computer, Playwright is a framework to do that. So, this project, they created something called a mobile right. Uh Jill Megidish, so this is his photo, and I believe it's in Germany. Uh primarily, it's a tool for mobile app testing and automation. Okay. Zero day clock. Uh thank you for sending me this this graph. It's absolutely amazing. So, it tracks the shrinking time uh from uh SVE, which is common vulnerabilities and exposures, uh disclosure to first confirmed exploitation. So, you see people found something and then uh they somebody exploited it. And it was the timing between was going down down down, and now it's actually negative. So, people haven't even publicly announced it, but hackers already found it and started explo- exploiting it even before the announcements. Yes, so this becomes harder and harder. Exploit survival also shift left year over year, showing attackers are moving faster than patch cycles. The practical takeaway is that patching alone is increasingly too slow. Exposure reduction and faster remediation matter more. Okay, Leopold Aschenbrenner. So, he's this famous financial guru who wrote situational awareness manifesto, uh who ran very successful hedge fund, and he overextended it and recently lost something like 30 billion out of 45 billion. So, went down to 10 billion. Uh so, he's now trying to calm the investors, and I don't know exactly what He's very smart guy. He used to work as a OpenAI researcher. And also, he got married in the middle of all that. Anyway, [snorts] Miko Say claims ChatGPT guided structuring Bitcoin based preferred stock helped him raise about 15 billion dollars. His core rule for ambitious builders, don't try to outwork automation, instead leverage AI as a force multiplier. He advises learning to ask AI for novel high-value solutions rather than mastering tasks AI can already perform. Uh yeah, so this is kind of unusual, asking AI for new ideas, for new solutions, not just automating mundane tasks. Uh, yeah, but you see it worked for him and worked very well. Okay, Meta compute. Meta has invested very heavily into the infrastructure and now they decided to rent out their GPU capacity. So now they're kind of competing with Amazon and other like providers of infrastructure. Uh, interesting. Okay, this is about jobs. August just started, not many layoffs yet. This is me as usual and thank you.