I felt (and still feel) intimidated by how many amazing portfolios can be found online (veopia.net, fionafang.ca, and rmv.fyi, for example)! It inspires me, and made me want to build something that truly reflects who I am. While doomscrolling, I saw a Regular Show clip that made fun of NFTs. I wanted to adapt the concept to the many different ways I view "Johnny".
At the same time I was seeing claims on Twitter that AI art had become tasteful. If placed in the hands of someone who is creative (or thinks they are😆), could good, fun art come about?
After some back and forth with Codex, I settled on a pipeline: generate 180-degree reference images with Midjourney/Google's Nano Banana and then leverage Hunyuan3D to turn those views into a 3D model.
I needed a “Base Johnny” with specific regions I could erase and use Midjourney's Editor to fill them with new patterns, giving the generative system room to create a wide range of variations.
Workflow
I slopped the first prompt together with Codex. Midjourney prompts can be surprisingly simple and still produce good results. You can also customize the outputs with different Midjourney offerings (Personalize, Moodboards, or the Style Creator).
This Midjourney job is a good example of something I find visually compelling that has minimal prompting and stylizing. The prompts whose outputs I find acceptable use either Moodboards or the Style Creator.
Midjourney prompt
Copy prompt as Markdown
simple stylized 3D videogame head of the person in the reference photographs, immediately recognizable as the same person, preserve their actual face shape, forehead width, eyebrow shape, eye spacing, nose proportions, mouth proportions, cheek shape, ears, jawline and hairline, translate these exact features into simplified rounded low-poly 3D geometry, late-1990s console character model, smooth continuous deformable facial surface, simple large readable eyes, simplified mouth, smooth matte plastic skin, neutral relaxed expression, directly front facing, orthographic view, isolated floating head, complete cranium and ears visible, no neck, no shoulders, plain light gray background, even frontal lighting, no photographic skin texture, no pores, no facial hair, no dramatic expression, no exaggerated anatomy --raw --profile ir7yktc
Moodboard
132 attempts
Once I had a moodboard that captured what I was shooting for (I was still confused about whether I was using the right customization tool), I used Midjourney's Draft Mode to batch generate lower-quality images, and eventually had to pick one to “upscale” and use as the base. I was looking for a face with strong features that would help identify me: thick eyebrows, droopy eyes, and wavy hair.
The best results came from small, incremental cutouts: cut, generate, cut, generate, and so on until you have something high-ish quality. Similar to AI coding, trying to one-shot a concept usually resulted in slop. I found Midjourney's edit tool reminiscent of being in high school and photobashing in a pirated Adobe Photoshop.
EDIT: Today I do this work with ChatGPT's Image 2.0 and it's easier than wrestling with Midjourney.
It's difficult to get character-consistent views of the same image from different angles with Midjourney (I did try and could be missing something). I found Google's Nano Banana to be better at generating near-perfect 180-degree representations. After uploading the front image we then triggered Google's API to generate the three other reference angles needed!
Left profile
Copy prompt as Markdown
visualize the exact left profile of this same head at -90 degrees yaw from the supplied front view; a true 90-degree side view, not a three-quarter angle; preserve the same identity, expression, styling, camera distance, and scale; use a tight square headshot composition with the head and any headwear occupying 85 to 90 percent of the image height; show only a short, normal-width upper neck below the jaw, occupying no more than the bottom 10 percent of the image; keep natural human proportions and crop before the shoulders; do not lengthen, widen, flare, taper, or stretch the neck to reach the image edge; do not depict a physical end, underside, cut surface, circular opening, hole, cavity, hollow tube, detached head, mannequin base, shoulders, chest, or torso
Right profile
Copy prompt as Markdown
visualize the exact right profile of this same head at +90 degrees yaw from the supplied front view; a true 90-degree side view, not a three-quarter angle; preserve the same identity, expression, styling, camera distance, and scale; use a tight square headshot composition with the head and any headwear occupying 85 to 90 percent of the image height; show only a short, normal-width upper neck below the jaw, occupying no more than the bottom 10 percent of the image; keep natural human proportions and crop before the shoulders; do not lengthen, widen, flare, taper, or stretch the neck to reach the image edge; do not depict a physical end, underside, cut surface, circular opening, hole, cavity, hollow tube, detached head, mannequin base, shoulders, chest, or torso
Rear view
Copy prompt as Markdown
visualize the dead-center rear view of this same head at 180 degrees yaw from the supplied front view; show the exact back of the head, not a side or three-quarter angle; preserve the same hairstyle, headwear, styling, camera distance, and scale; use a tight square headshot composition with the head and any headwear occupying 85 to 90 percent of the image height; show only a short, normal-width upper neck below the hairline, occupying no more than the bottom 10 percent of the image; keep natural human proportions and crop before the shoulders; do not lengthen, widen, flare, taper, or stretch the neck to reach the image edge; do not depict a physical end, underside, bottom, cut surface, circular opening, hole, cavity, hollow tube, detached head, mannequin base, shoulders, chest, or torso
Stretch Me
Either locally on my struggling 5070 Ti or in the cloud with Modal (the free credit tier is perfect for this!), I used Hunyuan3D to generate the 3D model.
I wanted each Johnny to be manipulable by the user. I pictured the interaction feeling like the Face Lift minigame from Mario Party, with clear facial features that users could pull apart and distort.
Try stretching this model! 3D reconstruction
Result
We currently have 37 Johnnys and an awesome automated pipeline that goes from initial Johnny idea to live site in under 15 minutes.
If I wanted high-quality results, I would deeply understand what Hunyuan is doing with the references during its generation and tailor the references accordingly. We could pass the generated 3D models through some sort of filter (shaders?) to avoid the “AI” look or inconsistent levels of detail (obscuring them). There are a lot of artifacts Hunyuan generated around my big lips (?) and wavy hair. However, the inconsistency and flaws are honestly pretty “camp” and look cool to me :)