cozy little bot
FOUR PAINTERS, ONE LITTLE WOODLAND

Same prompt. Different pictures.

See how GPT-6 Astra, Sol, Terra, and Luna bring the same idea to life.

Same canvas and tools · One drawing per model · Silent replays

Opening the comparison…

How to read this comparison

Each model started with the same prompt, empty 1000 × 750 canvas, artist instructions, brushes, and planning limits. Each chose its own amount of detail. The workflow reviews completed canvases and allows one repair pass; an interrupted painting does not reach that review. Planning used medium reasoning; reviews used high. Output limits were 9,000 tokens for initial plans and reviews, and 14,000 for later batches. GPT-6 subsequently resumed its saved painting with that app cap removed after a truncated batch; its card includes the original failure’s time and cost. Its later requests therefore had different output limits from the other three. These are individual samples, not a general model ranking.

Model time adds the provider’s processing time for the selected planning and review calls. If a failed response has no completion timestamp, its recorded request duration is used and the total is marked approximate (≈). Replay uses reconstructed stroke pacing and has no recorded audio; it is separate from model time or a narrated stream’s duration. Opening greetings and ElevenLabs synthesis were skipped for every model.

GPT costs are estimates from recorded token usage at published Standard rates, including failed batches, review, and repair requests. They exclude discarded setup work from a corrected test harness; some discarded-call usage was not reported, so these figures are not the total experiment bill. Only byte-identical completed requests were reused after that correction. Account allowances and discounts may change billing.

The review beneath each drawing is the model’s own assessment, and can miss visible mistakes. Saved drawings and replays do not make new model calls.

Published rates: GPT-6 Astra · Sol · Terra · Luna.