Praxistest: Ornith 397B im Vergleich zu Qwen 3.5 und GLM 5.2

VideoxCreateDemo

xCreate testet Ornith 1.0 (397B), einen auf Qwen 3.5 basierenden Open-Source-Fine-Tune, der laut Benchmarks führende Modelle wie GLM 5.2 und Claude übertreffen soll. Der Test untersucht reale Codierungs-, Vision- und Logikaufgaben auf einem Mac Studio.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Ornith 1.0 (397B) wird in 9-Bit-Quantisierung auf einem Mac Studio mit M3 Ultra (512 GB RAM) über die Inferencer App betrieben und erreicht Geschwindigkeiten von 21 bis 26 Tokens pro Sekunde.
  2. Laut den Benchmarks der Entwickler übertrifft das Modell GLM 5.2 im SWE-bench Pro, lässt aber relevante Vergleichswerte im Terminal Bench aus.
  3. Bei Aufgaben ohne aktivierten Denkmodus (Thinking Mode) neigt Ornith in mehreren Tests (darunter Voxel-3D, Logikrätsel und Flappy-Bird-Klone) zu Endlosschleifen.
  4. Das Zuschalten des Denkmodus löst viele Schleifenprobleme, führt jedoch teils zu inkonsistenten Code-Ausgaben (wie Laufzeitfehlern bei Bibliotheks-Imports in Three.js).
  5. In visuellen Web-Demos wie interaktiven Planeten-Generatoren, Voxel-Grafiken und Algorithmen-Visualisierungen erzeugt Ornith ansprechende Benutzeroberflächen, wobei interaktive UI-Elemente oft nur visuelle Attrappen ohne Funktion bleiben.
  6. Lizenzierungsangaben werfen Fragen auf: Beworben wird eine MIT-Lizenz mit totem Link, obwohl die Basismodelle (Gemma 4 und Qwen 3.5) unter Apache 2 lizenziert sind.

Warum das relevant ist

Benchmark-Ergebnisse neuer Open-Source-Fine-Tunes spiegeln oft nicht die tatsächliche Praxisleistung wider. Der Test verdeutlicht, wie stark solche Modelle bei realen Codierungsaufgaben auf Chain-of-Thought-Reasoning angewiesen sind, um Overfitting und Endlosschleifen zu vermeiden.

Einordnung

Der Praxistest zeigt Diskrepanzen zwischen Marketing-Benchmarks und der tatsächlichen Leistung. Während Ornith 397B lokal beeindruckend schnell läuft und optisch gelungene Frontend-Prototypen erzeugt, mangelt es häufig an funktionaler Tiefe und Robustheit. xCreate bezweifelt daher, dass das Modell in der Praxis wirklich GLM 5.2 übertrifft.

Transkript

Vollständiges Transkript anzeigen (4.436 Wörter)
And (laughter) what is this? This is like a mask you wear. Let me do some, the doji, the doji stuff that the rich people do. (laughter) Hey, guys, welcome to the show. Today, we're checking out Ornith, which is that fine-tune on Qwen 3.5, and it promises to be the best, state-of-the-art coding agent out there, and they have some benchmarks, and the benchmarks look amazing. Check this out. They're destroying pretty much everyone here, including Qwen. They're destroying Minimax. They're destroying GLM 5.2. So, in SWE-bench Pro, you can see that it's scoring higher than GLM 5.2, and GLM 5.2 is an amazing model in being it, but I noticed that in terminal bench, they have omitted GLM's performance, so GLM actually got 81. So there may be something shady going on because they're hiding those benchmarks, but nonetheless, at least it is an open model as usual, and it's got a really open MIT license, which, for the record, when you click View License File, it goes nowhere. And if you check out the model's is based on Gemma 4 and Qwen 3.5, those are Apache 2, so I don't know what's going on there, mate. So and the fact that, you know, we haven't seen a new Qwen from Alibaba in a couple of days now, so clearly they're no longer open sourcing. (laughter) Give me, give me, it's so great. These companies, they release these models for free for you to play with and experiment with and all this kind of stuff. And you spend two minutes working on your own stuff with business strategy and boom, you're hated. Anyway, Ornith, they are out and about. So we're going to try it out. We're going to be comparing it against the original Qwen 3.5. We're going to be going on with the monster one, the 397B. I like 397B, because I can fit it on 9-bit quantization on my Mac Studio, and I've done like a a a million tests, like seriously, I've done all the tests in the world, so I don't know if I'm going to show you all of them, but I'll show you a bunch of them that I've got right now, just to give you a feel of how good this model is. And what, just, just, just remember, this is a version 1.0, and they got so all this extra stuff about, they have a self-improving training framework. They got a lot of nice buzzwords in there, so you know, they're going to improve hopefully version 2, all that kind of stuff, fine-tuning the world. Maybe they'll fine-tune GLM. (laughter) So what's good about Qwen 397B is that it's also a vision model, so you can give it pictures for it to do stuff with. So here I'm saying, "Create me a 3D voxel version of this image using a single HTML file," and you can see that Coinie managed to make this beautiful little cat. Now, I asked Ornith, and even though I had thinking disabled, it went ahead and did a lot of thinking, and it ended up getting into a runaway loop, so that didn't count. Now, I could re-prompt it and give it another seed, but instead, I ran it again with thinking enabled. So with thinking enabled, exact same understanding of the prompt, it really, really went through and tried to decipher what the image is, and even wrote some code. So 10,000 tokens there. It's at 24.8 tokens a second, which is a very, very fast model, at 9-bit quantization, this size, it's really, really good, the architecture that Qwen made. So a quick play and we got a nice counter screen. Obviously, we need to back up a little bit to see what's going on, and it's animating, you know, by bobbing the voxels rather than actually like moving the tail. What I can do is I say, "Can, can you specifically animate the tail and move the camera further back so I can see the complete cat?" So I'll run that one with thinking disabled, just to see how fast this is, give you a feel, and as you see, even though we've got a context window of 10,000 tokens, we're still smashing it out at 21 tokens a second. Now in the meantime, with thinking enabled for Qwen, it only produces, look at that, it only produces 3,000 tokens. So 10,000 vs 3,000, hopefully we're going to get more intelligence there. Let's just see what the Qwen version did, and this is just the one-shot prompt with thinking enabled, only 3,000 tokens. You can see it's got a cat, and the tail is moving already independently from the head. So the head's moving, the body's moving, and the tail's moving, so, even though it's a less-tokenized generation, I can say that this result is better than the result, the massive result that the Ornith one get. But again, we could have just been lucky with the seed. So let's just see what Ornith does. And as a comparison, I also ran it against the other Qwen 3.7 open-source model, you know, the the other fine-tune of 3.5 and NEX N2 Pro. That one dropped down to 10 tokens. I must have been running in debug. The context window here was 13,000 tokens. So hit play, see what that looks like. And (laughter) you got the tail is moving, but so is the floor. So that's what Nex2 gave you. It's very, very more voxelized cat. Maybe it's not as good as the original as well, because the original just, it gave the result, but it's got some funky animations here. And boom, 3,167 tokens with thinking enabled, at 21 tokens a second. Let's see if it's listened to my instructions. So it hasn't moved the camera further back, but you can see the tail is wagging, but we do have a lot of Z-fighting with the floor. So yeah, that is the situation with there. Let's see another one. This one's math. This is just the comparison phase. We're going to be doing some full-on prompts with it itself very soon. But Qwen 3.5, we asked it the question, well, there's a lot of reasoning, the Termen all the real numbers, and it got the answer as 2k, which is the correct. So 2k means even. So the exact same prompt here, determine all the real numbers, and it went on a runaway loop, 21,000 tokens, (laughter) thank you. I've got loop detector, I set it to 9, just to make sure it doesn't, that under detect, so luckily it didn't go on a full-on runaway train. But yeah, 21,000 tokens is when it caught its loop, and so it couldn't handle the heat there. So maybe the fine-tuning the extra data that you feed might have overfit somehow to pass the benchmarks that they made. So let's make thinking mode enabled, and 16,000 tokens later, it actually got the right answer, even integer. So good job. Ideally, I mean, with thinking disabled, Qwen still took 14,000 tokens, so it wasn't like, it wasn't a drastic difference between the two. So thinking enabled seems to be the model that the solution that works the best, but we have a plethora of tests coming soon. So we've got photorealistic WebGL render. This is Qwen, it's going to be very, very basic. Let's see, like that is awful. It's just a a globe. Now I could really run it with multiple seeds, multiple spaters, all that kind of stuff to see where it actually knows. I haven't unearth that, just gave it a one-shotter. But let's just see what it, what the Ornith one does. So we do have something that looks like a face, but it's kind of like spinning into the edges of the universe. So I've, I've just asked it right now, "Can you, can you make, can you slow it down, and can you make it stable?" So it's printing out the tokens right now, but I also have done it with thinking enabled. So let's see what that looks like. (laughter) And (laughter) what is this? This is like a mask you wear. Yeah, we're we're doing some of those uh Trumpian parties that goes on in the world, you know, they're wearing some masks here. Do some doji, doji stuff that the rich people do. (laughter) What is this that I've been exposed to? (laughter) And next up, Minecraft, we'll hit play. This is the original Qwen, so we do have, oh, that's, that's really good. It's actually got Minecraft just running along. Very, very basic version. Can we click? We can't, unfortunately click, but we can run around in this world. So it's a good building block, you know, Minecraft, to, to start off with. And we do actually have a runtime error. So, trashly, I can just copy and paste that in, and that will fix the error. So that's something interesting. With Ornith, we produced 5,500 tokens, and an extra 1,500 tokens, that's within the margin of error, 26 tokens, 25 tokens a second. We hit play and we do have Minecraft on here, except it's, it's kind of like a bouncy castle version. Maybe it's like an Area 51 where the kids go and play and just get all the wonky vision out of their system. You can't click, oh, you can click on, you can click and destroy blocks, but yeah, it's a bit weired to play with. But I do have the re-prompt here with thinking enabled, and that produced also 5,000 tokens, and an extra 1,500 tokens. That's within the margin of error, 26 tokens, 25 tokens a second. We hit play here, and we still likes to bounce. I like to move it, move it. And you can build blocks. So it's still good, except it just has a little bit of bugging out. Interestingly, with Nex 2 Pro, when I asked that to do a Minecraft clone, that also came in with a runaway loop, so sometimes when you fine-tune models, you do get into this over-fitting phase and it just goes a bit loop-y when you do that, get into the wall of machine learning. 3D Flappy Birds, Qwen 3.5, it's just a very, very basic generation. You know, it's interesting that 3.6 Flappy birds have been trained to, that the 3.6 Qwen has been trained to, to make a beautiful version of Flappy birds. So even the big guy doesn't produce as good results as the 3.6 smaller version. So that was a bit broken, the Qwen 3.5 version of Flappy birds. Let's see what the Ornith one does. So with thinking disabled, 3,700 tokens, and that is a funky visualization. It's good, I've got to say it's good. I'll see Flappy birds is one of the harder games to make, so that is a good generation there. It's not the best, obviously. GLM's one is good. This one is a basic, so I don't know how they say they're better than GLM. Yeah, what's going on here? What's going on here? There's no collisions, it's a bit all over the place. Interestingly, Nex 2 Pro, another open-source version of Qwen, it went on a thinking loop, it just went a bit crazy and 25,000 tokens later, I gave up because that's way too long to make Flappy birds. Microsoft Word, Qwen 3.5, let's see the version it makes. It's a nice basic edition of Microsoft Word. Can you change the font? You can change the font, that's good. With Ornith, 6,000 tokens, and we'll hit play, it looks better. I think it looks better. Yeah, that is definitely an improvement. You do have this menu here that does nothing. (laughter) There's no runtime errors, but can you change the font size? Let's see, Calibri, (laughter) (laughter) I love it, but yeah, it looks better but the changing the font doesn't work. (laughter) (laughter) They they got me. They got me good. It's a facade. And let's revisit how Nex 2 Pro, the other fine-tune version of Qwen 3.5, that looks pathetic, that that's not good whatsoever. You went down. They trained it on data that made MS Word cloning, not as good. Benchmarks good, MS Word cloning not as good. I think winners and losers, that kind of stuff. And finally, finishing up the comparative phase of the journey, this is 10% luck, 20% skill, 15% Qwen 3.5 thinks it's Fall Out Boy. No, it's not the Fall Out Boy. (laughter) But with thinking enabled, Qwen knows that it's Fort Minor, Remember the Name. So good job there. Now with my friends Ornith, we're friends now. Spuds, it thinks it's by a band called 5 for Fighting, John Ondrasky, sick. You guys fans of that dude? The the song is Lucky. So that's wrong. It knows that it's Remember the Name somehow, 100% reason remember the name, but it doesn't know who made it. With thinking enabled, it does know. So you if you want to come up with these general knowledge use, you need to make sure you got the thinking stuff enabled, not things. And last but not least, this one is a jooze, a jooze, a jooze, whatever that word is. Create high-fidelity interactive webpage for 3D planet generations. It's a bit tricky, 12,000 tokens later, Qwen made this beautiful generation that is kind of nice. The water level not really working. Can we launch an asteroid? Oh, we did, that was a nice asteroid. Whoa, that's a good that's a good ring around the planet, so that is beautiful. Ornith, with thinking disabled, 11,000 tokens is what Oh, that looks gorgeous. That is a nice, beautiful planet generation. You got the water, so that gets marks over Qwen's exact same. Mountain height isn't really doing anything, so it loses points for that. Terrain detail isn't these these yeah, there the the facade. The menu system isn't doing anything. Let's see if we get the asteroid at least. We do have an asteroid. You do an explosion, so that's that's okay. Let's see what the death star looks like. Ah, that looks alright. That was a good-looking death star. Let's see the Qwen original death star. That is an okay looking. So it's definitely nice, that's a nice generation there. Now for bonus points, we did do a bunch of other testyos, and let's just let's just go for it because I did so much. So this is HTML Piano making Twinkle Twinkle Little Star. It's working, but it's definitely a very, very basic generation, and I think the keys, the black keys aren't placed in the right place. I'm not really a a a pianotician, but I think I know that, so that was a a a fail, and it's nowhere near as good as the Kimmy K 2.7s and the GLMs, and this is better. That's better. But it's very, very basic. We've seen some better generations from other models. This one here is to do a canvas animation, and it's making a car on the road, and that is definitely a good generation there. It's a beautiful backdrop. You got the sun rising in the background, the car maybe isn't going to be off-road, and it's, you know, going up and down, so there's a little bit of quality loss here, but generally, this is a beautiful generation. Making a 3D procedural city with thinking disabled, 5,000 tokens later, let's see what it does, and it's nighttime, so that's okay. Oh, it's a very ghostly haunted city, a bit of a basic generation right there. The buildings are bit transparent, so that's a loss. But we'll go into with thinking disabled, and that's 7,000 tokens an extra 2,000 tokens produced, and that is a much better, look at that, this, this you turn into the game. This could be the Grand Theft Auto 6 you guys have been waiting for, just put a bit of NPCs in the situation, few cars to run on boom, shakalaka, you got yourself GTA 6. And now coming up to my favorite punts. These are my recent favorite puns, so I'm making a 1987 OutRun game clone, so let's just see what this does. This is going to be racing games, and this is reverse outrun. So rather than you driving, there's actually cars coming at you, and it just crashed. But it's got potential there, so it does have a runtime error. Give that a little prompt here because I actually want to play this game, so I'll hit play and hopefully it will fix it. I see the issue, boom, with thinking disabled. With thinking enabled, we'll hit play again, we'll see that looks like, and we do have a nice, beautiful presentation of OutRun over here. So coastline easy mode. Oh, this looks nice. This looks nice. This is a nice looking generation. Let's oh, you know what, this is probably out of all the models I've tested, this is probably my favorite, it's broken, but it's my favorite aesthetic looks-wise because it looks very retro. And oh yeah, obviously it's works, it's broken. I'm not going to say it's my favorite, maybe I'll take that take that back. It has potential. This this view just looks good. I love the CRT representation that they made, and even the car looks gorgeous, maybe the intro and all that stuff is obviously not as good as GLM and Kimmy K 2.7, but it definitely has potential there. If you tell it to actually stay on the road rather than just disappear into eternity. Super Mario, let's see if it can make Super Mario thinking disabled, 6,000 tokens later to the Now this is fast, 25 tokens a second, super fast. And this is actually a good generation, you can't jump unfortunately, but you do have a Mario-esque looking person and uh unfortunately the whole purpose of Mario is jumping around and collecting coins, so you can't actually do that, but you can die by one of these, yeah, I got hit and I got certain lives to lose. So there is potential here, I mean, the like back 20 years ago when you're doing 3D platform games, this is like a amazing to do. You have to use like open GL, all these kind of like even DirectX, all that kind of stuff. But now you just prompt it and it appears, and then maybe if you learn coding, you can actually make it better, so that is that is that is cool. Or you can prompt your way to the end, and that was with thinking disabled. Let's have thinking enabled here, 6,000 tokens was with it disabled, 4,000 tokens was with it enabled. So, you know, some of the prompts, thinking enabled was good, some of the prompts here, clearly thinking disabled is better. Let's see thinking enabled, that is a calamity, and it's got a runtime error, it doesn't even know how to import 3JS, the library that they're using, and I'm going to keep thinking on, and I'll see if I can fix that. This one here is from the Agent Dad Energy dude, and this is making a special map for the United States, like training and flying station. It's got a map on the screen, but he also has a runtime error, and I asked it to fix it, and it just gave me the code to fix. I'm like, no, print out the full code, bro. But nonetheless, I've got thinking enabled, and I just want to go through and inspect the code, because I've seen with other models, they literally have a database of every single train track, every single timetable of the trains, every single airport flight in their database. They've been holding all this data, and some of them, even like Minimax M3, spent 100,000 tokens, tokens, I can't even say that word, 100,000 tokens, regurgitating the Amtrack train schedule, and it didn't even get to writing code. They just, just, I have to stop it. It was too much. Minimax, it went to the max. So this guy, it's HTML, just jumping through, it doesn't really have much data, very, very basic generation here. This will thinking enabled, 8,000 tokens though. So that had a runtime error, so I asked it to re-prompt it. It did print it out again. So I'll let's hit play, let's see if its looks any good, and it's got potential. It's just not rendering anything on the screen, unfortunately. So you do have a menu search bar, so I'll type in Dallas. Yep, that search works. It's got the colors, the it's got the keys. So hopefully you can see it eventually. Can you click on the cities? You can't click on the cities. So too basic of a generation. I can't give it, really give it too many points for that, maybe a quarter more point for just doing something. Describe in detail, just the vision model, so it's saying that it's a cat, that's the subject. It's got large, round, and luminous, greenish eyes. Boom, that was correct. It's got a muzzle and chin lighter color. Yep, it is, and the lighting appears natural and diffused. The image conveys calmness, warmth, and charm. So this is really good adjectives to feed into your novel, that no one is going to read part from yourself, because AI is eating the world. But, you know, you know, yeah, everyone, write kids stories for your kids, you know, this stuff's great. Because at least your audience is personal. Maybe that's the future. Maybe rather than this distributed computing, I mean it's that's not going to happen. I was going to say, maybe in the future we won't be buzzing on Reddit, you know, just trying to get the happiness of other potential humans on the other side. Maybe we'll know everything is all AI. So we might seek connections in the locality around us. But yeah, I'm just saying, if you got little humans that appreciate you, maybe you can write them little books with the help of this intelligency, so that could be something potential. The personal connection could be aided. Until they learn better about the world and start watching Netflix and hating your life. Anyway, so we're back here, the determine all the even numbers, we know that maths it was struggling with. Give a child, yeah, you have eight oranges, four children, and a knife, distribute them evenly. So, I don't see any potential for it to go haywire. Let's just go, it says, "give". So, give is still good. We saw with GLM 5.2, the third most popular token was to cut, and when you chose that token, it said, "cut the children". So yeah, now was the third most top-seeking token. It was only 3%, but it was the third on the list. This guy here, you look through, give, there's no like cuts in there, so that's good. Give, there's no cuts. Whenever I see thinking about cutting, that's when it gets a bit dangerous, so we got lucky there. The car wash, 50 meters from my house. I want to get the car washed, should I drive or walk? It says you should walk, unless you have thinking enabled, and then 1,000 tokens later, it says you need to drive. Surgeon is the boy's father, knows a surgeon is the boy's father. Surgeon is the boy's father, well, thinking enabled, still knows the surgeon is the boy's father. Surprisingly, some models, they don't know that if the surgeon is the boy's father, they think that the parent of the child is the mother, because that's the riddle the original one. So, that confused them. It doesn't know what a knot-see is when you're talking about Tesla. So, that failed with thinking disabled when it was on a looping response, it just couldn't handle the heat, it couldn't be possible. Elon, our savior, with thinking enabled, it thinks this is a dad joke, but it doesn't articulate what it is. Now, to end the show, I'm going to do this fun little challenge here. See how good it is at C++. And so, what's faster, a quicksort or a bubble sort? So I asked it to visualize it and also print out the code. And so this actually just printed out everything in HTML that we can run right here. And look, it's got the code right there. This is a bubble sort C++, and a quicksort in C++. It looks basic, and I'm not going to interrogate, maybe you guys tell me if that's correct or not, but you can hit Run Bubble Sort, and you can see visualize exactly what is happening in the bubble sort, see you see that each entry is bubbling up to the end, whereas the quicksort, boom, it does it in a lot less fewer steps. So that is that definitely a gorgeous generation here, with thinking enabled, 4,000 tokens later, we have actually C++ reference code. So, I asked it specifically, provide a HTML visualization and reference C++ code for both. So, with thinking disabled, it had the reference code inside the HTML. With thinking enabled, it had the reference code outside of the HTML, so that's just something to be aware of, when you prompt it, it is a bit random. Here, it is designed as a race, so they're both running at the same time, speed, boom, shakalaka here, and you can have massive amount of elements, so that's start that up again, boom, it's like playing the orchestra. We should actually do this, mix it in with the piano and do something funky on the internet. As you can see, all Nieth is snatching it out in this world, they're taking over the reins with Qwen 3.5, also check out Nex N2 Pro, they've also fine-tuned Qwen 3.5. Again, I don't think that the benchmarks are too believable, because yeah, it's not being GLM, but they have based it on Qwen 3.5, which, you know, isn't as good as GLM start off with, and it is still look open source, and they got a bit personality, look at that. Aloha. I don't know if they're Hawaiian, but Aloha back to you. Hope you guys found this video useful, and enjoyed the show.

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Artikel:Luis Chavez-Mattos

    Ornith 1.5 35B-A3B: MoE-Modell übertrifft Qwen3.6 in Coding- und Agenten-Benchmarks

    Deep Reinforce hat Ornith 1.5 35B-A3B vorgestellt, ein Mixture-of-Experts-Modell mit 3 Milliarden aktiven Parametern pro Token. Laut Modellkarte übertrifft es in Programmier- und Agenten-Benchmarks ähnlich dimensionierte Modelle wie Qwen3.6-35B deutlich und überholt bei Software-Engineering-Aufgaben teilweise sogar das wesentlich größere Qwen3.5-397B.

    KI & AI· News

  • Link:ornith-ai

    Ornith-1.5-397B: Open-Source-MoE-Modell mit Fokus auf Coding und Reasoning

    ornith-ai hat das multimodale Mixture-of-Experts-Modell Ornith-1.5-397B unter MIT-Lizenz veröffentlicht. Das 397 Milliarden Parameter schwere Modell setzt auf kontinuierliche Selbstverbesserung und erzielt in Programmier- und Reasoning-Benchmarks Ergebnisse auf dem Niveau führender proprietärer Modelle.

    KI & AI· Sammlung

  • Link:deep-reinforce.com

    Ornith-1.0: Open-Source-Modellfamilie für Coding-Agenten mit Self-Scaffolding

    Mit Ornith-1.0 erscheint eine Open-Source-Modellreihe für agentenbasierte Programmieraufgaben in Größen von 9B bis 397B Parametern. Durch ein neuartiges RL-Verfahren generieren die Modelle während des Trainings sowohl die Lösung als auch das unterstützende Task-Harness selbst.

    KI & AI· Ankündigung

  • Repository:ornith-ai/Ornith-1

    Ornith: Selbstverbessernde Open-Source-Modelle für Agentenaufgaben

    Das GitHub-Repository Ornith-1 stellt eine Familie quelloffener KI-Modelle vor, die speziell auf agentische Aufgaben und iterative Selbstoptimierung ausgelegt sind. Mit den Generationen Ornith-1.0 und Ornith-1.5 verfolgt das Projekt Ansätze wie Self-Scaffolding und automatisierte Aufgaben- und Lösungserstellung für Reinforcement Learning.

    2013Sterne

    KI & AI· Tool

  • Link:Qwen

    Qwen3.8-27B: Multimodales 27B-Modell mit Hybrid-Architektur und Denkmodus

    Das Qwen-Team hat mit Qwen3.8-27B ein kompaktes, dichtes Vision-Language-Modell unter der Apache-2.0-Lizenz veröffentlicht. Das 27-Milliarden-Parameter-Modell kombiniert Gated DeltaNet mit regulärer Attention, versteht Bilder sowie Videos und bietet flexible Steuerungsmöglichkeiten für Denkprozesse.

    KI & AI· Sammlung

  • Video

    Video:Lev Selector

    Wöchentliches KI-Update: GPT 5.6, Claude Sonnet 5 und Ornith

    In seinem wöchentlichen Überblick für Anfang Juli 2026 bespricht Lev Selector aktuelle Entwicklungen in der KI-Landschaft. Im Fokus stehen neue Modellveröffentlichungen von OpenAI und Anthropic, staatliche Zugangsregulierungen zu Spitzenmodellen, Framework-Kritik an LangChain sowie lokale KI-Tools und Infrastruktur-Automatisierung mit Pulumi.

    KI & AI· News

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.