Tessl und Snyk im Test: KI-Skills nach Anthropic-Richtlinien optimieren

VideoAI Native DevDemo

Simon Maple (Tessl) und Krzysztof Huszcza (Snyk) analysieren und optimieren ein produktives AI-Skill-Repository von Snyk. Mithilfe des Tessl-Agenten wird ein Skill auf Basis von Anthropics Best Practices bewertet, durch Refactoring verbessert und Sicherheitsaspekte wie Prompt Injection über Snyks Open-Source-Tool Agent Scan diskutiert.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Tessl bewertet eine bestehende `Skill.md`-Datei mittels LLM-as-Judge-Verfahren gegen Anthropic-Empfehlungen; der Ausgangswert für den Snyk-API-Target-Skill lag bei 87 %.
  2. Positiv bewertet wurden klare Pfadführung (Routing) und die Abgrenzung zu verwandten Skills, kritisiert wurde hingegen die zu hohe Dichte und fehlende Conciseness in der Hauptdatei.
  3. Der Befehl `tessl review fix` lagerte vier Codeblöcke für Authentifizierungsmethoden in separate Referenzdateien aus, um Progressive Disclosure zu nutzen und Context Bloat zu vermeiden.
  4. Durch die Auslagerung stieg die Bewertung von 87 % auf 90 %.
  5. Snyk setzt intern darauf, persönliche Entwickler-Skills schrittweise zu standardisieren, insbesondere für sensible Vorgänge wie Deployments durch Plattform-Teams.
  6. Für Sicherheitsprüfungen von Skills (z. B. auf Prompt Injections) verweist Krzysztof Huszcza auf das frei verfügbare Tool 'Agent Scan' im Snyk-GitHub-Repository.

Warum das relevant ist

AI-Skills und Coding-Agenten werden zunehmend für Entwicklungs- und Betriebsabläufe eingesetzt. Schlecht strukturierte Skills überlasten das Kontextfenster von Sprachmodellen, während unsichere Prompt-Formulierungen Risiken wie Prompt Injections bergen. Strukturierte Prüfroutinen und Progressive Disclosure stellen sicher, dass Skills sowohl token-effizient als auch sicherheitskonform bleiben.

Einordnung

Die Demonstration veranschaulicht eine wachsende Disziplin im AI Engineering: das systematische Qualitätsmanagement von Instruktionsdateien (Skills). Anstatt Prompts und Kontextdateien unstrukturiert wachsen zu lassen, zeigt das Beispiel die Modularisierung nach dem Prinzip der 'Progressive Disclosure' – Details werden erst bei Bedarf nachgeladen. Bemerkenswert ist zudem die Unterscheidung zwischen funktionaler Skill-Qualität (Tessl) und sicherheitsbezogener Überprüfung gegen feindliche Prompts (Snyk Agent Scan), was auf eine beginnende Aufgabenteilung in Entwickler-Toolchains hinweist.

Transkript

Vollständiges Transkript anzeigen (3.530 Wörter)
Simon Maple: Okay, so we have the results back. It’s moved it 3% from 87 to 90. Krzysztof Huszcza: Hey, now it’s a combination of everybody is just writing their own skills, right? Skills are a little bit personal, right? Simon Maple: You prompt your LLM, like your your coding agent, like you do a session, and at the end of the session, you want to have this repeatable thing, right? So you say okay, just save it. Save it as a skill, right? Simon Maple: So I was walking through the Expo Hall at AI Engineer in San Francisco, and I saw this brilliant pure light and I was blinded by it. And I looked up and it was Krish walking towards me, and he said, "Simon, I need you to evaluate a skill of mine." And I said, "Krish, what skills have you got?" And you said, "I have no idea." Simon Maple: So, so what we're going to do, this is Krish. Krish, what do you do at Snyk? Krzysztof Huszcza: I, uh, head product strategy at Snyk. Simon Maple: Product strategy at Snyk. Krzysztof Huszcza: That is, indeed, the case. Simon Maple: That sounds very important. Krzysztof Huszcza: That's, maybe. Simon Maple: I would have thought like there would that need to find someone smarter than you to do that role, though. Krzysztof Huszcza: Uh, Simon left, unfortunately. Simon Maple: Oh, I did. I did. Now there's no more smart people at Snyk. Krzysztof Huszcza: Okay, okay, so you were the last one, you were the last one. Simon Maple: But but you're a good alternative. That's fine. That's fine. So, Krish, uh, what we're going to do is we're going to go to GitHub. We're going to go to Snyk. Krzysztof Huszcza: Yep. Simon Maple: Uh, we're going to, let's search for a skill for a skill.md. Let's find us – oh, and not MSD. Let's search for a skill.md. Krzysztof Huszcza: Yep. Simon Maple: Let's find something that we're going to – we're going to uh we're going to evaluate and and run a review across. Krzysztof Huszcza: Okay. Simon Maple: Uh. Krzysztof Huszcza: Just find this, this raw MCP skill. Simon Maple: This one? You know this one? Krzysztof Huszcza: Yeah, this is for creating targets. So it's our dynamic, kind of scanning testing, security testing product, and it's – it's API-first, so we created these skills to help our customers add targets to to the product. Simon Maple: Okay. Krzysztof Huszcza: Using Cloud. Simon Maple: So, let's come here, let's make a directory called Snyk skills. Uh, then I'm going to, I'm going to git clone this fellow here. Uh, and then we'll cd into saw-mcp and I'm going to run Tessl agent in here. Krzysztof Huszcza: It's very exciting. Simon Maple: Oh, that's not the agent. Krzysztof Huszcza: I know. Simon Maple: Look at that. See? Okay. Um, right, so we're going to run Tessl agent here. Now I'm going to say I I'm going to say, okay, let's, uh let's, let's say please, please. We don't know how long we're going to be above agents for, so it's important it's important in the runup that we apply it. Just in case, right? Krzysztof Huszcza: No, I'm doing the same thing. I'm doing the same thing, just in case they take over, right? Just to the end. Simon Maple: And they got, they got good memories. Krzysztof Huszcza: They got, they got, they got good, they got good memory. Simon Maple: Let's say, "Please take a look in this project. In fact, let's have a look in this project and and tell me which skills exist." Let's have a look at the Let's have a look at this, just that one or if there's uh multiple. Simon Maple: I'm going to Yolo this. There we go. Krzysztof Huszcza: So that is your agent. That is the agent that you build. Simon Maple: This is the Tessl agent, so it's an agent which essentially is, uh Oh look, we've got two there. It's an agent that essentially is there ready. I'd say it's dedicated to making sure your your agent enablement is as good as it can be. Looking at optimizing the skills and the the environment that you have from a context point of view. It also works a little bit more to get you to the next stages, small step by small step, into into closer to a software factory. Awesome. Okay, so we have two here. We have the we have the raw – we have the saw – the saw target the saw API target configuration. Uh, no, we have – it's the same thing. Krzysztof Huszcza: No, so one is for API targets. So this is basically for testing APIs and testing web applications. So one of – one of them is kind of configuring the the test for the web target, so like if you have like an application to test. Simon Maple: Yeah. Krzysztof Huszcza: Like a web URL, right? So that's that's that's one skill. And the other skill is just like if you have an API to test and you want to send kind of like dynamically test the API for security. Uh, so that's the other skill. So there are two skills here. Simon Maple: So which one do you want to – which one should we evaluate? Krzysztof Huszcza: Let's do the API one. Let's do – Simon Maple: API target configuration. So API target configuration skill. Uh we'll say, "Use the Simon workspace." Simon Maple: Right. So what this is going to do is, it's going to run, hopefully it will run Tessl review run. And what Tessl review run does, in fact, it's looking at Tessl review help there, uh, and then it's going to right there it is. It's run Tessl review run. Krzysztof Huszcza: Okay. Simon Maple: It's running it against the saw API target configuration. And what this is going to do is, it's going to run a review which is an agentic review. So it kicks off uh a review, LLM as a judge style. It looks at the uh It looks at the best practices from Anthropic, and it will say, "How close, how, how well is this skill been written," compared to the the Anthropic best practices. Now that's not to say that the Anthropic best practices are absolutely perfect. Krzysztof Huszcza: Okay. Simon Maple: So if you get 0% here, it might just be that the Anthropic best practices are rubbish. Krzysztof Huszcza: Okay, maybe, maybe, maybe. Simon Maple: I'm using that as a caveat, just because we don't know what's going to come back here, so I'm using that as a caveat. Krzysztof Huszcza: Okay. Simon Maple: Whatever the score is, we should be able to do a fix, and we should be able to say, "How can we improve this a little bit?" And so we'll have a look at it and see, you know, where we can potentially improve. Krzysztof Huszcza: Oh, hopefully it's not 100 then, because then there's nothing to improve. Simon Maple: We'll see, we'll see. I'll tell you what, if it's 100, Krzysztof Huszcza: Okay, you give – Simon Maple: What what do you what do you want to bet? If it's 100, you can have this Tessl swag, Dilly Dallying. Krzysztof Huszcza: Okay, okay, so – Simon Maple: I like your hat. How about we swap? Krzysztof Huszcza: Okay, so if it's – if it's not a 100, I'll take your hat. Simon Maple: Okay. Krzysztof Huszcza: It is 100, you can have this – Simon Maple: Okay, I'm, I mean, I mean, I mean a deal. Krzysztof Huszcza: Deal. You shouldn't have said deal, because I've just seen the score. Simon Maple: I, I saw the score as well. Yeah, okay. So, let let's take a look. So the score is actually, do you know what, 87%? Krzysztof Huszcza: 87! Wow! That's pretty good. That's pretty good. Simon Maple: Now let's have a look at where uh So it's already a strong usable skill. Sorry, let me just let me just grab this. Sorry. Krzysztof Huszcza: My goodness, the hat as well. Simon Maple: There we go. You've got this hat and I'll I'll wear this one. Krzysztof Huszcza: I mean, that's a good trade. That's a good trade. Simon Maple: This is already Well, your head's Your head's pretty big. Your head's pretty big, Krish. I'll take I'll take that off for for many reasons, but, um I'll put it there for now. This is already a strong usable skill. The reviewer liked that it has clear routing, API target only, not generic web target creation. Good sibling-skill disambiguation. Ah, now this is super important. Right? So you've got two skills and it and it's important to be able to determine when to use one, when to use the other. And this is very often when we talk about activation, one of the biggest problems of uh of agent activation isn't just "do I use your skill or not," it's "do I use the right skill when I've got a combination of multiple skills to pick from." So that's actually really useful. Simon Maple: Good, handle the duplicate target warning. Strong actionable authentication examples. That's great. That's great. But the skill is too dense. Too much detailed authentication reference in line with the Skill.md, making it less concise and less progressively disclosed. So what that means is, a skill doesn't get fully loaded normally. It gets progressively disclosed. So you load pieces of a skill at a time. So what it's probably going to suggest is you take a lot of that Skill.md, break it into sub-files and reference those different sub-files. And then when the agent realizes, "I want to go deeper into here," it will load that. But if it doesn't need to, it won't load that. So it avoids context bloat. Krzysztof Huszcza: Right, so this is basically this progressive disclosure to avoid context bloat. Simon Maple: Exactly. And I can tell you if you look here, the progressive disclosure is 2 out of 3, conciseness, probably from the size of the skill. Simon Maple: Specificity, one of the hardest words to say in the English language. Krzysztof Huszcza: I cannot – I cannot pronounce this. Simon Maple: No. No. No, no one actually knows what that means. This is the true story. No one, no one knows. Many, many have studied it. But but nobody knows. So, we have to just say it's 2 out of 3. But we don't know how to fix it, because we don't know what it means. Krzysztof Huszcza: All right. Simon Maple: There's a bunch of recommended fixes. What we're going to do, because we're agentic developers, is we're going to ignore it. And we're just going to say, "Run Tessl review fix." And we're just going to assume that the fixes are good and we'll apply those fixes and we'll blindly follow it. Krzysztof Huszcza: So, basically, I'm trusting Tessl to to to do that and just push to production. Simon Maple: Yeah. Yeah. Perfect, man. I'm I'm all in. I'm all in. I'm all in on the AI like just just let let's go. Let's just go. Simon Maple: Well, what we're going to do is we will show a diff afterwards, and we'll see what it's changed. It's not quite like that. We'll see what it's changed and then we can accept or, or you know, reject. Krzysztof Huszcza: Can you have an agent to see the diff, so, I mean like I just want to go. Simon Maple: Exactly. That's exactly what we'll do. We'll get the agent to do a git diff. Oh yeah. What, you mean getting an agent to review it? Krzysztof Huszcza: Yeah, exactly, like one agent is just doing the fix, the other can just do the diff review, right? And now we can go for a beer. Simon Maple: You had me at "let's go for a beer," so that's fine. Simon Maple: So, so let's have a look. Uh, where was that? Can you remember, Snyk API web maybe? Krzysztof Huszcza: I don't remember, man. Simon Maple: Config, it was config, skills, there we go, and it was the API target and it was, oh, there's a single Skill.md. This is the problem. Krzysztof Huszcza: Yeah, it's just bloated like, it's too – Simon Maple: This is bloated. So, so what I suspect it's going to do is, it's going to reduce this and you'll start, you'll see, uh extra files here. Uh that's what I suspect. But we'll see, because what happens is it tends to read too much of this, too much of this. It might update the description as well. We'll see because very often the name and description the two pieces that the agent reads before knowing when to activate the full skill or not. So we'll see what happens. Simon Maple: So, how do you how do you folks use, obviously, thanks to Snyk as well, because we use Snyk within Tessl to test every single skill that we have on the registry from a security point of view. That's true. But how do you use Tessl, how do you use skills internally? Krzysztof Huszcza: Internally? So, I think right now it's a combination of everybody is just writing their own skills, right? Skills are a little bit personal, right? So like, you know, like you do like you you you prompt your your your your LLM, like your your coding agent, like you do a session, and at the end of the session, you want to have this repeatable thing, right? So you say okay, just save it, save it as a skill, right? It's kind of you have your personal skills. That's kind of one thing that we do. The other thing that we do, we try to slowly, slowly standardize uh across the organization, so kind of share these skills with other people, right? And also standardize on some skills. So, for example, if there is a skill that knows how to push stuff to production, we want to have like that part of our platform team publishes a skill that developers can use because then it can pass all of the security and quality kind of requirements, right? So, like if our developers kind of write a skill to push to production themselves, it's like a little bit of a fre- freaking out our security team. So like we're doing this to two things broadly, right? Like internally how we use skills. Simon Maple: Yeah. Simon Maple: Let's talk about actually, there was the uh, what was the report that you that you did? Uh it was our toxic skills. Toxic skills, yeah. Toxic skills. And it's it's amazing to see how many skills, through, kind of just like poor poor wording, poor writing, actually introduce potential security issues, prompt injection, those types of things. What can people do differently to actually write skills that are closer to uh uh you know, having fewer, or no, security issues in there. Krzysztof Huszcza: I mean, use Tessl, right? This is is – use Tessl agent is one uh, yeah, yeah. Simon Maple: I tell you what, if Tessl gets it to 100%, I keep the hat. If it's not, we swap that. Krzysztof Huszcza: We swap back, no. I don't want to swap back, though, you want that one. You like that one, okay. Okay, okay, okay, I want to keep the hat. Simon Maple: Okay. I mean, use Tessl, you say you sneak, we have – Simon Maple: So you can you can run the Snyk tool independently though, right? You can run it out the out the command line. Krzysztof Huszcza: We have we have a tool called Agent Scan, so that that tool is available on GitHub. GitHub/Snyk/Agent-scan, I think is the is the is the repository there. And you can just go go there, run the tool, uh point it at your already know, it kind of scans your file system, but you can also point it at a at a skill, and it will give you like a security kind of advisory review of of of of that skill. Right? So that's for the skills that you write yourself, but also very important for third-party skills because kind of like these third-party skills might be malicious. Right. Right? So, uh, yeah, so so that's that's recommended for sure. Simon Maple: Yeah. Krzysztof Huszcza: Uh, yeah, so um, I think that's that's the that's the main thing really for kind of skill security. Krzysztof Huszcza: There will be a different categories of of problems for first-party skills and third-party skills. So I think this space is still evolving, so we still need to understand, you know, what what advice to give to developers who are developing their own skills on on kind of security, right? That's that's that's that's maybe a bigger problem to solve then just kind of looking for prompt injections, which are malicious, that's that's that's a little bit easier. Yeah. Like working progress, but we'll get there, we'll get there. Yeah, yeah, cool. Cool. Simon Maple: Okay, so, uh we have the results back. So it's only moved it plus 3%. I'm still going to wear this then. It moved it 3% from 87 to 90, and it reached that it reached the target score of 90. Let's have a look. Acting on the judges' uh content judges' suggestions, I've moved four detailed API authentication method code blocks out of Skill.md into the new references. Ah, so it has moved that. Krzysztof Huszcza: Right, so so it's split the skill into smaller pieces. Simon Maple: Absolutely. So now the the one of the big changes is, it will now progressively disclose that to the agent as and when it needs those specific pieces. Uh so that's the kind of big one, and it's there's a whole bunch of uh there's a whole bunch of changes here that you'll that you'll actually see. So first a, have a look at this now. Now we've got this references here, and you'll see extra hosts, logout detection, sequence formatting, and you've also got the Skill.md, and at this stage, we'll almost certainly – Krzysztof Huszcza: Yeah, it's much smaller, yeah. Simon Maple: Uh Hang on, is it this one or No, sorry it's this one. This one, yeah. This is – this one's the one I wanted. It's got the API authentication MD and in here, uh you will at some point reach out to that, uh – hasn't as much as I thought. I was probably expecting it to be to be more, just to fit the full for the full code for the authentication method, uh go and have a look at this. Right. And so it reached out, and maybe it will maybe if we was to run a couple more times, it will probably actually, maybe you know, do do that again. Krzysztof Huszcza: We see how the 10% to go, right? So. Simon Maple: Yeah, yeah. And that's the thing, right? Because sometimes, we can then have a look at what what else it wants to improve and see if it does want to improve and iterate across. But uh, hey, that's not too bad. Not too bad. Krzysztof Huszcza: That's That's great. That's great. That's great help. Simon Maple: Yeah. I'm going to use that now. Krzysztof Huszcza: Yeah. Especially as I need to share my skills, you know? I don't know if they are any good. So, like I would like, you know, like if it if it did my personal skills, like maybe that's fine, right? I want to share with my team, you know? I want to I want to use something like that. Simon Maple: And you can and you can and you can add this into you know right into the into the uh into the pull request checks into so you can as soon as a pull request happens when your Skill.md, you can rerun the review, make sure it's not regressing and those types of things. And it's available for free for developers. Absolutely available for free, so you can you can you can try it out, you can get a you can get an account straight away and you have a number of uh credits on your free on your kind of like free account and then you can just use that throughout the month and it gets re- replenished every month, so. Krzysztof Huszcza: Amazing. Amazing. Amazing. Simon Maple: Awesome. Krish, absolute pleasure. Thank you so much for joining us. Then, and there's you can have that 3% for free, and take that away with you. Yeah. Krzysztof Huszcza: Thank you very much. Thanks a lot. Simon Maple: Thanks for the help as well. Appreciate it. Krzysztof Huszcza: Yeah, cheers. Simon Maple: Cheers.

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Artikel:Anthropic

    Claude Code mit Skills erweitern: Architektur und Best Practices

    Anthropic dokumentiert das Skill-System für Claude Code. Entwickler können Arbeitsabläufe, Workflows und wiederkehrende Prompts über SKILL.md-Dateien definieren, die auf dem offenen Agent-Skills-Standard basieren und dynamische Kontext-Injektion sowie Subagenten unterstützen.

    KI & AI· Tool

  • Repository:anthropics/skills

    Anthropic Skills: Referenz-Repository und Spezifikation für Claude Agent Skills

    Anthropic stellt in einem öffentlichen Repository Beispiel-Skills, Templates und die Spezifikation des Agent-Skills-Standards für Claude bereit. Skills bündeln Anweisungen, Skripte und Ressourcen in Verzeichnissen mit einer zentralen SKILL.md-Datei, um Claude wiederholbare Aufgaben wie Dokumentenerstellung, Code-Tests oder Workflows beizubringen.

    174.341SternePython

    KI & AI· Sammlung

  • Artikel:Barry Zhang, Keith Lazuka, Mahesh Murag

    Anthropic stellt Agent Skills vor: Modulare Ordner für spezialisierte KI-Agenten

    Anthropic führt mit Agent Skills einen offenen Standard ein, um KI-Agenten wie Claude modular mit prozeduralem Fachwissen und Skripten auszustatten. Skills bestehen aus Verzeichnissen mit einer zentralen Konfigurationsdatei und ergänzendem Code, die vom Agenten dynamisch und bedarfsgerecht geladen werden.

    KI & AI· Ankündigung

  • Link:Anthropic

    Agent Skills: Offener Standard für modulare KI-Agenten-Fähigkeiten

    Agent Skills ist ein offenes Format zur Erweiterung von KI-Agenten um domänenspezifisches Wissen und Workflows. Der Standard basiert auf Ordnerstrukturen mit einer SKILL.md-Datei und nutzt Progressive Disclosure, um den Kontextbedarf gering zu halten.

    KI & AI· Sammlung

  • Video

    Video:Hyperautomation Labs

    Wie Anthropic-Teams Claude Skills einsetzen und wie man sie selbst baut

    Mitarbeiter von Anthropic nutzen intern standardisierte Claude Skills statt manueller Prompts für wiederkehrende Aufgaben. Das Video erklärt anhand von Praxisbeispielen aus der Finanz- und Rechtsabteilung des Unternehmens, wie das Konzept der Progressive Disclosure funktioniert und wie man eigene Skills in Claude Code oder claude.ai anlegt.

    KI & AI· Anleitung

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.