Show HN: A Dalle-3 and GPT4-Vision feedback loop
A simple web app that loops DALL·E 3 image generation with GPT‑4 Vision’s image-to-text prompts is prompting people to explore how visual ideas “evolve” over multiple iterations, often drifting into surreal, cosmic, or steampunk themes. Commenters treat it like an automated version of party games such as Telestrations or exquisite corpse, experimenting with custom prompts, language translation, and constraints to see which elements stay stable and which quickly distort. Alongside the creative fun, there is scrutiny of OpenAI’s safety filters, rate limits, and diversity directives, as well as concern about pasting API keys into third‑party tools and frustration with the official ChatGPT image UI.
Concept & Overall Reactions
- Tool chains DALL·E 3 and GPT‑4 Vision: each image is described by GPT‑4V, turned into a new prompt, then re‑rendered by DALL·E, iterating into a “telephone game” of images.
- Many comments express delight and amusement; people share surreal chains (corgis, gnomes, goats, AI painting itself, Earth cycles, etc.).
- Some find it creatively inspiring; others say outputs are noisy, repetitive, or “cheesy,” reinforcing existing doubts about AI art.
Prompt Design & Evolution Patterns
- Custom “meta‑prompts” matter a lot: e.g., “replace everything with corgis,” “make it more whimsical,” “increase intensity each time,” “make it weirder,” or “describe via a specific lens.”
- Longer, hyper‑detailed descriptions make images more self‑similar across iterations; compressing to very short text increases drift.
- Several note convergence toward certain tropes: cosmic/space, mushrooms, steampunk, futuristic cityscapes, psychedelic posters.
- Multi‑subject prompts often collapse toward a single dominant subject, often the first mentioned.
Model Behavior & “Understanding”
- Some see impressive visual coherence (correct shadows, reflections, 3D‑like structure) and stable themes.
- Others argue GPT‑4V often misdescribes images in ways no human would, suggesting pattern‑matching rather than real “understanding.”
- There’s debate whether apparent misalignment is due to model limits vs. inherent lossy compression from image → short text.
Safety, Bias & Content Filters
- Users frequently hit safety rejections and opaque “rate limit” or “unsupported image” errors; behavior is described as inconsistent and sometimes misleading.
- Shared internal prompt text for DALL·E shows enforced diversity in depictions of people (race/gender mixing), which some applaud and others find intrusive when they want culturally specific scenes.
API, Costs & UI / Feature Requests
- Tool is front‑end heavy; users appreciate it as a faster, less frustrating alternative to OpenAI’s default UI (which is criticized as slow, buggy, and hiding prompts).
- People discuss rate limits, tiers, and per‑iteration costs; 10‑image runs cost only cents but add up.
- Requested features: forking chains, graph of lineage, starting from an uploaded image, single‑use or capped keys, better handling of 1‑image/min limits, and fixing share‑link/“keep going” bugs.
Security & Emotional Reactions
- Strong reluctance to paste API keys into third‑party sites; some advise temporary keys and revocation, others remain uneasy.
- A few report a visceral sense of repulsion or uncanny “wrongness” from AI art, even when finding the experiment intellectually interesting.