Vollständiges Transkript anzeigen (2.044 Wörter)
Okay, so over the past week or so, I've been using GPT-6, and it's increasingly clear that code in some sense is going away, or at least in the way that we used to think of it. This model in particular, it is exceptional in a number of different areas. One in particular, it has a really good sensibility in terms of how to gather different context, but also the design sensibilities. And the agentic capabilities are some of the best that I've used before.
Now, I'm a huge fan of Anthropic and the Claude series of models, but this model in particular, I've really been focused on using this one over the past week, and the results are really impressive. So in this video, I want to do an applied practical demonstration on just how easy it is to create beautiful applications now and leverage a number of different tools within this.
What I'm going to do within here is I'm simply going to describe what I want to envision. I'm going to say, I want to create a yoga studio website, very similar to something like the Lululemon website. And what I want to do is I want to leverage and create images leveraging the Higgsfield CLI, and I want to generate models with different articles of clothing for eight tiles on the homepage, and I also want to be able to click through to each of those models and the respective articles of clothing and see the product page for all of them.
Now, if you have used these coding agents before, you probably have a general sense just in terms of what they're capable of doing. But increasingly, what I've found with GPT-6 Astra in particular is there are surprisingly little gimmicks that you have to hand it to make it work. And what I mean by that is there are a lot of tools that are built into it, or you can just specify tools that you might have locally, and it will know how to leverage those things without you having to write these intricate skill files or having to have a really robust agent.md, or all of these different tools that previous generations of models really did benefit from. This is one of those models where you can simply just describe what you want in a lot of contexts. You don't need to focus on trying to encourage the model to loop through and continue going like these Ralph Wiggin patterns that we saw earlier in the year. You can simply tell it a task and it will go off and perform it for you.
And the one thing that I want to call out with this model is it does feel really good at some of these tasks that you can spin off. And oftentimes, the tasks that I ask of GPT-6 Astra can go 10-20 minutes where I can just hand it a task and then I can move on to another task. It's one of those models where if you have got comfortable with the idea of parallelization and managing different agents, it does really good at that. What we'll notice within here is when it does need to spawn up sub-agents, it will do exactly that. Within here, we can see that it generated all the particular assets for us, and we can drill down and take a look at what it generated for us. We have all of these different articles of clothing and the models like we had wanted for our yoga studio application or website. And then it can actually go through and create the site for us.
The one thing that I wanted to show you within this video is just some novel ways in terms of how easy you can actually make applications come to life that aren't just these text-in, text-out type of inputs. I think there are a ton of use cases where you can leverage images and videos in a ton of creative ways, and I'm going to be diving into that as soon as this is ready. And then here we go. So as you can see within here, there are some subtle tweaks that we can make like we can move the model down, so the top of her head isn't being cut off. But at the bottom here, you can see all of these different tiles like I had asked for. We even have these beautiful little hover effects where I can go and I can click through, I can open up the page, and this was just with one prompt. I didn't specify anything about the design in particular to make it look like this, but it really has incredible design sensibilities. Next, what I'm going to do is I'm going to say, "Now what I want to do is I want to take the photo of the model in the hero area and then what I want to do is I want to pass that photo into Higgsfield and generate with Seedance, one of their new Seedance models, I want to generate a photo of the model doing a yoga pose. Have it within a yoga studio, and let's have that studio be based with a background of New York City and the skyline behind her. So, if you have..."
Now one thing that I want to mention, so within Codex directly, you do have access to also generate images, which is something that I don't think a lot of people actually realize. While I'm showing you with the Higgsfield CLI, you do have the option where you can generate images directly from the Codex or ChatGPT desktop app and have them weave in those different assets within the web app that you're building. But additionally, the thing that I love with Higgsfield is ever since Sora has been discontinued, is what you can do is you can leverage a ton of very powerful capabilities to really make your website come to life. This can be things whether it's videos, images, leveraging a bunch of different models that aren't just within OpenAI, but there are also some very powerful capabilities for rendering 3D assets, which I'll show you in an upcoming video as well. As you'll see within here, again, just with a sentence, I'm just talking to my computer, it's able to go through, find exactly what I described, it's able to run the particular command for the Higgsfield CLI, and it will go through and leverage the Seedance model like I specified.
And then within here, you can see I prepared a Warrior 2 pose in a sunlit studio with a view overlooking Manhattan. So on and so forth. It's going to go ahead and proceed. It also has the ability where you can answer questions where you aren't forced or blocked to answer questions, but you can have a little bit of time where if you do wanted to specify particular inputs, there is this subtle aspect that will come up where if it does draw your attention, you can add in additional input. So here is the video that it had generated for us. So what I'm going to do within here is I'm going to say, "Now let's add in this video to the hero area of the website."
And then here we go. That video that we had just generated is now a part of our website. So, again, just to emphasize how easy it is to make really incredible things nowadays. Okay, so next in this video, I partnered with Higgsfield to showcase their Genjutsu model. And what this model allows you to do is you can pass in any video with reference artifacts, such as different images, and be able to translate what you want to have over different scenes. Just to give you an idea in terms of how you can leverage this. So it's incredibly straightforward. What you're going to be able to do is if you have a reference video, this could even be a video of you recording yourself or having another AI model generate what you ask of it, and then you can pass in that reference video and translate it into a ton of creative areas.
In the case of what we're building, like in the context of a yoga studio, there might be certain approved poses that we want to transpose with the particular models across all of the different product pages for instance. And then the cool thing with this is if you have solid reference videos, this could be something like riding a bike within New York City, or this could be, like in the case of the yoga studio, you could change the background depending on where people are accessing the website from. I've been playing around with this with a ton of different videos, and what we can do within here, just to show you a few. Within here, here's an example of a model within a yoga studio and as you can see behind here, it's within the setting of London. And if you wanted to change out just the background, you can see in some other videos that I had generated, you can have this be in other scenes, but still have it be the same reference video. We have the same pose, but the background here is within New York. Instead of actually having a crew record different models within different pieces of clothing, which could be prohibitively expensive, you could imagine having a reference video of what you want to generate and then within here I can select the prompt and I can say, "I want to swap in the model for the reference image that I have and have it keep the motion transfer."
And then the really cool thing with Higgsfield is because they have a CLI, they have an MCP, but then they also have this really rich web app, it makes it really easy to generate a whole host of different, whether it's images, videos, or audio. And that makes it something really powerful in combination with Codex because famously, ChatGPT did discontinue Sora in favor of their text generation models and all of these agentic capabilities that they're really focused on. So having a good outlet for being able to leverage all of these different capabilities and all of these different state-of-the-art models that are coming out from different companies, whether it's Google, or Seedance, or Minimaxs, or what have you, you have an option where you can tie all of this together and leverage it within one platform itself.
Here is the video that it had generated for us. Now the great thing with this is what we're going to be able to do is I can simply go over to Codex and I can say, "Let's leverage the Higgsfield CLI and let's pull in the video assets of what we had just generated with the Genjutsu model. Also I want to add in the image for what we had generated for this and weave it into what we had just built." And so overall, it just gives you a really good sense just in terms of how you can leverage this. This can be used obviously in the context of a ton of really creative elements. You could be making a music video, or you could use this in a ton of video editing type of task, but you could also imagine leveraging videos in the context of building an application. Maybe you actually want to build an application that shows different yoga moves, or it's an exercise application. Being able to have something like this with the consistency and being able to potentially configure the different character or different model in different applications, this just gives you a really easy way in terms of how you can do something like that. So overall, I just wanted to do a really quick one showing you GPT-6 Astra in combination with what you can do with some of these models from Higgsfield, but otherwise, if you found this video useful, please like, comment, share, and subscribe. Otherwise, until the next one.