kw: experiments, art generation, artificial intelligence, simulated intelligence, clock dials
It takes a child about six years to learn to read an analog clock. It takes even longer to learn to draw one correctly. Art generation software seems to be on the verge of reaching this milestone.
This is the winning image in the experiment I describe below.The prompt read
A beige wall in a store selling clocks, showing four circular wall clocks, each reading exactly 3:27.
There are two parts to showing the correct time on an analog clock dial. Firstly, the minute hand has to point to the correct minute mark. Here, all four clock dials show this correctly.
Secondly, the hour hand must point to the correct location. In this case, the hour hands should all point 27/60ths (45%) of the distance between the 3 and the 4. I made careful measurements on the original image. The four hour hands point identically to a spot just above the second of the marks between 3 and 4 (~38%). They should point just a little below it. They all are where they should be at 3:23, about six minutes too early. There is one other possible anomaly in the image. The fifth clock shown, near lower left, reads 10:10, as nearly all advertisements for clocks have for decades. This abundance of training material causes clocks reading 10:10 (or sometimes 1:50) to dominate images of clocks produced by nearly all art generators.
This image is the best result that an art generation model has achieved since I began using DALL-E2 almost four years ago. This model is named GPT Image 2. I believe it is the model used when you ask ChatGPT 5.4 to create an image. At this point, GPT earns an A- from me.
For this experiment I tested ten art engines, GPT Image 2 and nine others. I had each of them produce two images and selected the best one to be shown below. My focus was the most recent models offered by Leonardo AI and OpenArt. There is a lot of overlap between these two "umbrella" sites. I began with Leonardo AI, using GPT Image 2 first (The other image by GPT was nearly this good, but the four clocks all pointed at 3:25, and all four hour hands were pointed the same as these). Now I'll discuss the other models tested, in order.
Google's flagship model is Nano Banana 2. It excels in realistic imagery. In this case, it set the beige wall I prompted as though it were above a pass-through into the workshop of the clock store, nicely blurred for a bokeh effect.The times are not all the same, and none of them is 3:27. They range, according to the minute hands, from 3:09 to 3:18. None of the hour hands is correctly aimed, but a couple of them are close.
I had noticed earlier than when I prompt Nano Banana 2 for an image including a clock, and specify a time to show, the result is only approximately correct. However, I've usually accepted that because for the scene I am creating, almost any time other than 10:10 will do. So for other reasons, I use Nano Banana 2 at least as frequently as GPT Image 2.
The Ideogram model is a private brand. It was among the first to be able to correctly render text (most of the time). This version is P-Image Ideogram.Text: good if I had asked for any, based on prior experience. Clock dials: not so good. All show 10:10, and even more so, the hour hands point almost exactly at the 10, where they should be one-sixth of the way from the 10 to the 11.
The dial designs are interesting. The blue and black ones are quirky, mixing genres in a way few clock designers would do, and the blue one in particular has inaccurate Roman numerals.
The Seedream series is owned by ByteDance; this is version 5.0.The minute hands are exactly right. The hour hands have all gotten ahead of themselves, pointing almost all the way to the 4, as though the time were 3:57 rather than 3:27.
Although the other clocks shown in the background are bokeh-blurred, I can see that each reads a different time. This is a plus for Seedream, something I can take advantage of when I actually need multiple clocks in an image, such as depicting a clock repair business.
Finally, under the Leonardo AI umbrella, Lucid Origin is a Leonardo "house brand." It is perhaps a little dated, but new enough to be quite good with many imaging tasks.As far as showing the time, although these clocks don't show 10:10, the model has entirely gone off the rails. None of the four pseudo-times shown is close to 3:27, and none of the hour hands is in a plausible location.
This "good-looking but implausible" attribute is characteristic of older art generation models, not just regarding clock dials.
Now we'll switch to OpenArt. GPT Image 2 and Seedream and some others are also available here, but OpenArt seems to specialize in rather off-the-wall specialty art models. The first shown here, which is also the newest, is Recraft V4, a specialty art model.Recraft is also among the first to have good text rendering. It is also known for creativity; to a certain extent, one needs to be very specific with one's prompts, and it will still usually throw in unexpected elements. For that reason, I use this one when I'm creating more whimsical images.
The times shown are all over the place, similarly to Lucid Origin, but are generally more plausible, as though Recraft has a better handle on the relationship between the minute hand and the hour hand. But check out the second dial, with 22 up there!
WanAI, part of Alibaba, produces the Wan family; this is Wan 2.7. They focus more on video generation, which requires better consistency from image to image. Come to think of it, it has been a good while since I worked with a model that allowed me to set the seed. Most models now keep the seed hidden.This is also a model that will throw in extra stuff. The four central clocks all read 10:10, and all have the hour hand pointing right at the 10. The other clocks nearby have a variety of readings, and sometimes an extra hand or two. They are more "clock-ish", not really clocks. That makes this a model I might use to create a scene on another planet, where the idea of a "clock" only roughly approximates our own.
Grok Imagine Image 2.0 belongs to SpaceXAI. It aims to be good-to-excellent with everything.Here, it's better than most of the others, but falls behind GPT Image 2. The four clock dials all have their minute hands aimed at :25, but the hour hands are stuck at the 3.
I do like the little motto at far right. I use this model when I'm trying out a prompt on several models, looking for creative variations.
The Chinese company Kuaishou Technology has the Kling art models; this one is Kling 3 Omni.Here, the times shown appear random, and there is no relationship between the hour hands and the second hands; none is plausible.
Rather than the "beige wall", this appears to be a display board. The rest of the clock shop is a nice touch. This is another of the models I use to see what creative stuff it might throw in.
Here is the last experimental image. Alibaba also owns the Qwen art models; this is Qwen Image 3.0.The times shown are approximations to 3:27. None has the hour hand in a proper location. Close, but no cigar.
Qwen models had an early focus on more accurate people, when DALL-E3 was still putting extra fingers and toes and other distortions on them.
This is the only image of the ten that shows prices for the clocks. It must be channeling a future time; the prices are about twice as high as similar clocks at most stores...at least the ones I shop at!
Clocks are one of the edge cases I use, firstly to judge the accuracy and comprehension of an art model, and secondly to detect a deepfake. In particular, I have seen a few AI-generated podcasts recently that showed the avatar in an office setting with a clock in the background. The clocks' hands didn't move.
For the interested: another "tell" is that AI generated images and videos in particular are sharper than "real" photography or videography. As shown above, creating bokeh isn't hard anymore. But the avatars are a little too well-defined.










No comments:
Post a Comment