TL;DR

I turned a photo of my laptop in the car and a MacBook meme into a ten-second looping film through a conversation with Codex.

  • Process: I described the idea, shots, and transitions. The GPT model in Codex worked out the prompts, code, and editing tools to make them happen.
  • Images: OpenAI image generation in Codex created or edited six scene images from my supplied references.
  • Video: MiniMax H3 Max Turbo on fal generated six five-second clips, using a starting and ending image for each move.
  • Style: a close-up becomes a wide reveal, the SUV freezes while UK scenery pans past, then the film returns to the exact opening screenshot. Codex added the reply animation, meme, timing, and sound locally.
  • Video API spend: an estimated US$0.30, excluding image generation and Codex usage.

I posted a photo of my laptop setup inside my car. One MacBook, a small grey table, a mattress on the left, and window covers. The caption was: “Added a new photo to my battlestation 🖥️”.

Kitze replied: “ELABORATE anson”.

Later, over lunch, I watched a random video about video editing and the different shots and camera angles a videographer should know. It made me think I could try learning it myself. Kitze's comment gave me something to practise with: I could make a video reply and see whether I could bring the scene in my head to life.

That became the starting point for a ten-second film. I wanted to begin writing a reply, delete it, and let the photo do the explaining. The camera would enter the image, pull back to reveal the black SUV outside a gym, then take the same little workspace around the UK before returning to the original post.

I supplied the idea and reference images, pointed out what looked wrong, and kept refining the sequence. The GPT model running in Codex figured out how to implement those requests: writing the generation prompts, choosing the editing tools, and writing and running the code. My part was describing what I wanted to see and giving feedback on the result.

All six scene images were generated or edited using OpenAI’s image generation in Codex (ImageGen), including the UK locations. My original photo, post screenshot, and the meme were the supplied references.

The moving shots were generated with MiniMax H3 Max Turbo through fal, then Codex assembled the finished edit locally. My estimated spend on the video API was US$0.30.

The original post showing my laptop inside the car and Kitze asking me to elaborate

The other reference was the MacBook bell-curve meme. That’s Kitze’s actual setup at the top. Apparently one MacBook only counts if it has an audience of other screens. At either end of the curve, the answer is “rawdog the macbook”. That was the connection to my post: I was just using my laptop in the car.

My initial request was to make a few frames where the image “comes to reality”. Start inside the workspace, zoom out to the whole SUV, then show it in different places.

I wanted the changing viewpoint to do the explaining. The close-up would make the laptop feel like an ordinary desk setup. Pulling back would reveal where it actually was. From the open boot, the camera would show the arrangement inside the car, then move out to a wide shot with the whole SUV and the gym in view.

Those directions became a close-up of the laptop, a rear view through the hatch, and a wider rear three-quarter view showing both the open boot and the side of the SUV. Each shot revealed something the previous one couldn't show. That was the part I wanted to practise: using the framing and camera movement to tell the story.

The boot view: upright seat on the right, mattress on a raised frame with storage on the left

Codex used image generation to create a sequence of reference frames: the interior, the open boot, the whole SUV outside the gym, Cornwall, the Lake District, and the Scottish Highlands. The locations beyond the original photo were generated scenes.

I asked for a ten-second script with a freeze and a pan for each place. Once the whole SUV was revealed, I wanted its angle and framing to stay consistent so the changing surroundings would carry the journey. Codex translated that into a script where the car held still while the scenery swept from right to left, then everything froze at the next location. It gave each transition 0.30 seconds and the Highlands a slightly longer final hold.

I wanted to review the whole sequence with its images, so I asked Codex to put it in a standalone story folder with an index.html and an assets directory. The page showed each shot, its timing, camera direction, sound, and transition. Being able to see those decisions beside the frames made the next changes much more concrete.

The opening and ending then became more specific. I supplied the exact screenshot and asked for the film to start and finish there. At the beginning, the cursor clicks Reply, types something, then removes everything. Instead of sending an explanation, we enter the image.

Codex used “Well, basically…” as the draft and built the reply interaction as an animation over the screenshot. Nothing gets posted.

The meme found its place between deleting the text and entering the real scene. It appears as an unsent image attachment. The camera pushes toward the right-hand monk, the “rawdog the macbook” caption, and the MacBook, then cuts to my actual laptop setup.

The MacBook bell-curve meme used as the unsent image reply

That gave the film its final sequence: attempt a written reply, delete it, try the meme, enter the laptop, reveal the car, travel through the UK, and return to the post. It fitted into eleven scripted beats across ten seconds.

For the moving shots, I specifically asked to use MiniMax H3 Max Turbo’s image-to-video endpoint on fal. Codex used a starting image and an ending image for each generation, with a prompt describing the move between them.

Codex generated six clips through the API:

Clip Movement
1 Laptop interior → open boot
2 Open boot → full SUV outside the gym
3 Gym → Cornwall
4 Cornwall → Lake District
5 Lake District → Highlands
6 Highlands → laptop interior

For the API requests, Codex set each clip to five seconds at 768P, with balanced prompt expansion. That produced thirty seconds of source footage for it to assemble into the ten-second edit.

GPT turned my descriptions of the car's layout into details in the generation prompts: the upright rear-seat backrest on the right, the raised mattress frame and storage on the left, and the table forward in the rear-seat area. It also worked out a camera path over the seat’s shoulder for the reveal, so the seat could obscure the laptop naturally.

I wanted the SUV to stay still while the scenery changed. GPT put that direction into the travel prompts and worked out an additional editing step: using a fixed cutout of the approved car over the moving backgrounds. That helped keep the car consistent when the generated clips varied around its edges and wheels.

The compositing needed its own cleanup. Early versions showed doubled edges and blur around the SUV. Codex refined the car outline and adjusted the background blur so it did not smear the generated vehicle into a dark trail behind the fixed foreground.

The SUV in the generated Scottish Highlands setting

For the edit, GPT chose Python, Pillow, OpenCV, and FFmpeg and wrote the code to combine the six MiniMax clips with the reply animation. It kept the screenshot as a background and animated the typing, deletion, and meme over it. These were technical choices it worked out from the sequence I described. Codex also created the clicks, typing sounds, whooshes, and quiet air locally, with no voiceover.

The return needed to close the loop properly. After the Highlands, the camera moves back toward the laptop, and the interior recedes into the photo inside the original post. The unsent reply and meme disappear before the interface comes back into view.

To meet my request for the same opening and ending frame, Codex restored the opening screenshot by 9.60 seconds and held it to the end. It also checked the exported video’s first and last decoded frames and confirmed they matched pixel for pixel. Its final export was a ten-second MP4: 300 frames at 30 fps, at 1280 × 720.

My estimated spend on the MiniMax H3 Max Turbo video API was US$0.30: six five-second clips at fal's promotional rate of $0.01 per generated second checked on 3 September 2026. That covers thirty seconds of generated footage, edited down to ten seconds. It excludes the earlier image generation and Codex usage. The amount is based on our recorded requests and the listed rate; we could not verify the final invoice because the API key did not have permission to read billing.

The story page keeps the six original clips alongside the finished film, images, script, and download links. I also asked for a Storyboards gallery on my portfolio so each story has a thumbnail that opens its full page.

I wanted to start picking up video generation, and this gave me a small project to learn with. I got to practise the shots and angles I'd watched over lunch, describe the movement I imagined, and see how the generated clips came together in an edit.