Showing posts with label prompts. Show all posts
Showing posts with label prompts. Show all posts

Sunday, April 20, 2025

My Flying People experiment

 kw: generated art, ai experiments, surveys, simulated intelligence, prompts, prompt adherence

Introduction

This is quite long so I include headings.

I remembered a science fiction story that I read decades ago, about a planet of creatures that looked a lot like humans, but had wings and flew about. Pondering a way to illustrate the central idea of the story, after some experimentation, I came up with this prompt:

On a planet with low gravity and dense air, many winged men and winged women are flying into and out of very tall buildings with doorways and landing platforms at every level

I was primarily interested in the great variety of imaging Models offered by Leonardo AI, and I had thousands of credits available. I surveyed nearly all the possibilities in Classic State, which yielded 184 images, all based on a particular Seed; more on that anon. I also used the prompt to produce much smaller suites of images with Dall-E3, DesignStudio, Gemini, and ImageFX. To jump ahead to a useful conclusion, I found that Dall-E3 produced images closest to the meaning of the prompt; in the lingo of the field, it has the greatest Prompt Adherence. Here are two examples, the best of the 24 images I gathered from Dall-E3:



I am showing two images because the first has the best people and the second has better buildings. When prompting an art generating program, it takes persistence and cleverness to get good adherence to a complex prompt. BTW, did you notice the figure with wings on backward?

Models and Styles

Leonardo AI (henceforth Leo) has thirteen Models or Presets and each has a number of Styles, as many as 23. This is a list of the 13, with the number of Styles each includes, and the cost in credits per four images, the usual number. The free program assigns 150 credits daily. It is easy to use them up pretty fast! The basic monthly subscription assigns 8,500 credits per month for $12. These pile up fast unless you are very active. I wish there were an in-between subscription level such as 3,000 credits for $5.

The total number of Styles is 61, and the number of Model+Style combinations is 194. I did not quite use them all. Several Styles are named Monochrome or include B&W in their name and those don't interest me.

Seeds

Leo also has the option of using a fixed Seed. Leo's help text says that this Seed is used to generate the "noise" that the reverse-diffusion process starts with to produce an image. That is not all there is to it, because the Seed value also influences the Transformer (the routine that figures out object identity and placement, which defines the goal toward which the diffusion algorithm operates). Seeds in Leo (also DreamStudio) have six digits.

When I set a fixed Seed, I tend to pick Seed values that are a multiple of 999,999/7 or 999,999/13, all of which have six digits except 999,999/13, which is 76923. After a period of noodling around, I found the region around 428571 (3*999,999/7) most interesting and settled on 428575. All the images shown here were produced using this Seed.

Selection 1: Leonardo Lightning Style Gallery

No matter how I may group the images, it would be tedious for the reader to wade through a discourse on 184 images. Thus I picked certain themes. I chose primarily certain Styles across all the Models, but I will start by showing the gamut of Styles for the Model called Leonardo Lightning (or LLightning), which runs faster than the others and costs the least—two credits per Medium-size image (1280x720), while for most Models the cost is 3 or 4, and it goes up to 12, and even higher where larger images are available in certain Models.

Forthwith, the gamut of LLightning images, screen-captured three across from File Explorer, so the file names can be seen:




The last image above is from the next set, from the Model Phoenix 0.9.

LLightning is unique among Leo's Models; 16 of the 21 Styles shown here produced things with wings, but only 7 are winged people. The third Style, Cinematic, has what appear to be human-sized bats, but they may be humans with bat wings. It is hard to tell which, even on the full size image. Two others show beetle-like flyers, two have birds, and the flyers in the other 4 images are unidentifiable. Finally, the images from None and Unprocessed are apparently identical, which I find logical: both claim to be doing nothing "extra" to the Model. This is also seen for the Model Cinematic Kino, the only other Model that offers both None and Unprocessed.

Compared to the other Models, this mix in LLightning is interesting. Four Models (the two Phoenix versions and the two Flux versions) adhere to the "winged" part of the prompt 100%, although only the two Phoenix versions have winged persons in all images, while the two Flux versions have more birds and fewer people. On the other hand, two Models (Graphic Design and Stock Photography) never produced wings on anything. Most of the Models yielded low percentages of winged persons. Anime is unique in a different way. Half of its Styles produced images with winged humans and two Styles have airborne humans without wings (one hopes they are floating, not falling). Only two Anime Styles had no wings at all.

Selection 2: Model Galleries for nine Styles

Style 1 = None (10 uses)

Now I will focus on particular Styles as developed in different Models. The Styles to be presented are those having larger numbers of Models that use them, in descending order by usage numbers. The first is None, meaning no Style was applied. This (non-)Style is available for the largest number of Models (10), with a modification to be mentioned below. I used a search to isolate each set of images, and a quirk of the search function is that the images are presented in reverse order.

In the next-to-last row of images, the first two appear identical. Upon very close inspection I find a few tiny differences. In the row above that, the second and third image are, so far as I can tell, identical. This shows that behind the scenes Portrait Perfect and Cinematic Kino use the same engine, as do Graphic Design and Illustrative Albedo. So in this case the 10 Models produce 8 unique images (discounting a few nearly invisible differences in one case). These images show what each Model produces when it is not constrained by a Style.

Style 2 = Dynamic (8 uses)

Dynamic is the default Style for the 8 Models that use it.



For the Dynamic Style, the image I like best is for Phoenix 0.9. It is the best match to the image I had in mind after reading the story, so long ago. Comparing all these images with the prior set, I find that for Flux Schnell the Dynamic Style produces the same image as None. Similarly for Flux Dev, Phoenix 1.0 and Phoenix 0.9. For the other Models, these Styles produce significantly different results. Three of the Models—Illustrative Albedo, Cinematic Kino and Lifelike Vision—have what I call "winged structures", although one of them (for Cinematic Kino) looks like an immense beetle.

At this point note that "doors at all levels with landing platforms" is seldom found, and that is primarily in Phoenix 0.9.

Style 3 = Portrait (8 uses)



As we'll see later for the Fashion Style, Portrait often emphasizes a central figure, although in the case of Illustrative Albedo, that figure is a flying structure, looking like a giant crab with 6 wings. The image from Lifelike Vision has a figure with wings that are more like hang glider wings, rather than bird wings. But, wings they are.

Style 4 = Stock Photo (7 uses)



Stock Photo is the only Style used by the Stock Photo Model. Its image is almost identical to the one produced by Cinematic Kino, but not entirely (one must look hard to find the differences). The other 5 Models all yielded winged things with this Style, but only the image for Phoenix 0.9 has winged people.

Style 5 = Ray Traced (7 uses)


Now we can start to see that certain Models, such as the two Phoenix Models and the two Flux Models, have the primary influence. In other cases, the Style seems to be "stronger" than the Model. Other than having a brighter and more colorful appearance, Ray Traced is similar to Stock Photo. For this Style, only the Phoenix Models produced flying people.

Style 6 = Illustration (7 uses plus "Anime Illustration")



It is more clear with these, compared to the prior sets, that the Style is paramount for LLightning, Illustrative Albedo, and Lifelike Vision. Here, Anime Illustration Style in the Anime Model joined the Phoenix Models in producing flying people.

Style 7 = Fashion (7 uses)


Here is a Style that produces winged people, almost 100%! Lifelike Vision's person has a winglike, flowing robe. Perhaps the presence of about 8 wings on the person in the Cinematic Kino image makes up for that… As one might expect from the name, a "fashion model" is front and center.

Style 8 = Creative (uses = 7)



Following trends we've seen above, some of these differ from other Styles for a particular Model, while a couple of them are more similar. Only Flux Dev and the Phoenix Models produced flying people. The flying things in the LLightning image seem to be huge insects though one of them has feathered bird wings. I don't know why a scattering of these images feature hot air balloons.

Style 9 = 3D Render (7 uses, with two added sizes)



As mentioned earlier, most of these images were generated for the Medium 16:9 size, which is 1280x720 for most Models, but is 1184x672 for the Phoenix Models and the Lux Models. I did a side experiment to show the effect of changing image size with Phoenix 0.9. You can see in the file names for "LPhoenix09" the numbers 01a, 01b, and 01c. The image sizes are 1184x672, 1376x768, and 1472x832. When I want to use any Phoenix image for a 16:9 wallpaper, these sizes are a little off. Leonardo can upscale an image so it is larger than 1920x1080 (full HD), but then trimming is needed to get an exact ratio. Changing the size causes changes in the overall image, though the three images are similar.

For this Style, the Phoenix Models produced flying people, but the others differed: LLightning and Flux Dev made flying beetles, Flux Schnell made birds, while Illustrative Albedo and Lifelike Vision produced lots of balloons but nothing with wings.

Wrap-Up

The Leonardo AI Models definitely have minds of their own. Only images produced by Phoenix 0.9 had both elements, flying people and towering buildings with landing pads at upper levels. Few other Models had the buildings as I had asked. Generally, the images are somewhat inspired by the prompt, sometimes rather distantly. It would take a lot more investigation, letting the Seed be randomly produced, to see whether any particular Model+Style is capable of better compliance.

I use art generating software as a "commissioned artist". A few of the Models in Leonardo, the more costly ones, are reasonably compliant. Programs other than Leonardo vary in their Prompt Adherence, and Dall-E3 is probably the best. The rest of Leo's Models, in most of their Styles, are fun to experiment with, but are unlikely to yield results that well match anything but the simplest queries.

I did not test a rather new way of using Leonardo AI, Flow State. It has a different way of doing things. Neither did I turn on "AI Enhancement" for my prompt in those Models that offer it. What we have here is complex enough already.

The folders of images I produced are a useful Gazetteer of possibilities that I can use in the future to select an image generation routine.

Monday, February 03, 2025

AI can't tell time

 kw: ai experiments, prompts, ai art, failures

In various corners of the universe of knowledge, SI (Simulated Intelligence) is manifestly ignorant. This can be seen in certain everyday tasks, such as reading (or drawing) an analog clock. I heard mention that most advertising for clocks and watches shows the hands set to 10:10 because ad writers think that has the most attractive appearance. Since SI has no knowledge of what clocks even are, or how the hands show time, they are dependent on their training image sets, which are only useful if there is text accompanying the images. I decided to see how various art generators would handle this prompt:

An image of a very decorated mantel clock showing the time as 4:15

I first used Gemini. This is the result, showing my prompt (one has to tell Gemini this is to be an image or picture):

The program did a good job with the decoration. That is its strength. But, sure enough, the clock's hands are pointed at 10:10, or very nearly so. If you look closely, the hour hand is exactly at the 10, where it should be 1/6th of the way to the 11. 

Gemini produces only one image at a time, in contrast to all the other programs at my disposal.

I next tried DreamStudio, the most recent program I use. I set the number of images to make at 2, because I pay for credits, and each image costs something. Using the same prompt:


DreamStudio is playing a trick in the first image. The hands have a "head" at both ends, so the time being indicated is ambiguous, but one interpretation is still 10:10. Though the hands have different shapes, it's also hard to tell hour from minute hand, so eight interpretations are possible! Don't try to teach your kid to read a clock that has such pathological hands!

The second image at least has a hand pointed at the 4, but given that the other is pointed at the 8, indicating 4:40, the hour hand should be a bit more than halfway between the 4 and the 5.

Next victim: ImageFX (driving Imagen 3, the same as Gemini). It is free to use, so I let it run four images. I also left the aspect ratio at 16:9, the setting I usually use with this program.


All hands point to 10:10. the expected result. Next, Leonardo, using the "bare bones" Leonardo Lightning style and the default (Dynamic) substyle:


Here we find an interesting variety of responses. 

  • Upper left: The hands are so nearly the same length it's hard to say if this is 10:10 or 2:50, although whichever hand is the hour hand, it's pointed right at the digit, not advanced as it should be.
  • Upper right: This looks the most like a real clock. The hour hand is between the 4 and the 5. It still isn't showing 4:15.
  • Lower left: The hour hand is near the 6, but on the wrong side of it, unless it is just a little too far over and the time should be read as 5:40.
  • Lower right: 4:40, with a misplaced hour hand, as seen before.

Finally, here is the response from Dall-E3 in Bing:


Assuming I've figured out correctly which hand is which in each case, the times shown are 10:07, 2:50, 10:09 and 12:55.

So there you have it. Not one 4:15 in the bunch.

Thursday, August 31, 2023

Automated abstract art

 kw: experiments, simulated intelligence, art, generated art, images, prompts, abstract paintings

I wanted more pieces of generated art to use for Zoom backgrounds. I happened to see an abstract mural at a restaurant so I took a picture and looked it up. The artist was Lyubov Popova (1889-1924). I downloaded a couple of Popova's pictures and one by Zaza Tuschmalischvili (b. 1960), cropped portions of them, and in three sessions, uploaded each portion to Dall-E2 as a "seed" for some outpainting. The prompt for each session was the same with the exception of the artist's name:

Angular abstract painting in the style of [artist]

Here are the three results.


With the third one, by a living artist, I worked back and forth, eventually erasing nearly all the original seed, so it became more of an inspiration than an integral part. I like all three, though they are rather garish for use as backgrounds. For that purpose I lightened them by adding about 50% brightness to each, with this result:


Depending on the audience, one of these could be an appropriate background. I tried to create abstract paintings without specifying style; here is one of the results, based on the prompt

Abstract painting with vertical grain, melancholy color scheme, low key


The location of the color bar shows that I painted this in vertical orientation, then turned it sideways. Even lightened up, it seems inferior to the ones above, as a background at least.

For the sake of academic meetings or among those who love libraries, I also ran the prompt

A wall of bookshelves filled with books of all colors and many sizes

I outpainted it to two sizes. The smaller one, which is close to HD, I also lightened for use as a background. The larger one (with more shelves) is nearly 4K.


On an HD monitor these all fill the screen. On a larger one, the images with fewer pixels may not do so. This image competes in my affections with the "Big Library" image, shown in this post.

Wednesday, August 30, 2023

Automated art, illustrating Outpainting with DALL-E2

 kw: experiments, simulated intelligence, art, generated art, images, prompts, photo essays

A charity, prospecting for contributions, included some Christmas cards with their appeal. This one really appealed to me. I decided to scan it and upload it to Dall-E2 so I could extend it by Outpainting. In earlier posts on automated art I wrote of Expanding, but Outpainting is the main term used by OpenAI. The original painting is "Holiday Social", copyright Geno Peoples 2017. The artist's website shows this and many other paintings in this "luminous cottage" style and other similar styles (note that the website is not secure).

I tried scanning at various resolutions with descreening, both automated and manual. What worked best was a cell phone photo, manually descreened and reduced in pixel resolution. I suppose I could have downloaded it from the website, where an image 1100x717 is available. If you want one of his paintings big enough to print on canvas you can buy it there. The image above is 1024x763, and is a bit narrower than the original.

When you log into Dall-E2, you can either start with a prompt to generate an image or upload something you want to edit. I prepared this image, and a slightly larger rendition (1450x1080), to experiment with uploading and outpainting. Eventually I ran three outpainting sessions. Grammatical note: In a lengthy discussion like this, I tire of circuitous locutions using the pronoun "one", so I will instead use "you" language.

When you open a Dall-E2 session and upload an image, you have the option to crop it to a square or leave it uncropped. In all cases I skipped cropping. If you crop the image, it is squared up to a size of 1024x1024 pixels or smaller. Every Generate action is done within a Generation Frame of size 1024x1024.

First Session 

I first used the 1450x1080 version, which was loaded at full size when I clicked "Skip cropping". It is necessary to add a prompt before you can click Generate. I began by adding material at all four corners. The image below shows the Edit screen after the first two additions. The buttons at the bottom are, from left to right, Select (for moving the Generation Frame), Scroll (for moving the entire image, Frame and all), Erase, Add Frame (for initiating a Generation Frame), and Upload (which I haven't used. I suspect it lets you add an object to the image but I don't know). The Generate button is above this part of the screen, on the right side. There is a download button next to it, and I use it frequently, keeping in mind the admonition in the gray box at lower left (you may need to click on the image to see it large enough to read that text). Downloaded images are PNG files, about 3-4 times the size of JPG files, but with no compression losses. 

The prompt at this point reads "A holiday visit on a snowy day at sunset in the style of Geno Peoples"; I changed it later.

When you click Generate at any point, you have four choices to choose among. If none is acceptable, you can click Cancel and Generate again (consuming a credit...sigh). After outpainting in all four directions I had this result:

The two extensions in the sky added versions of the sun near sunset, which I didn't like. I combined more outpainting with editing. The next image shows the upper right corner during this process.




The white circle at center right is the eraser. I got rid of the bright bit of sky and an unusual looking window in the tall building. I also removed the sun image at left (not shown in this crop) during the following Generate event.

After getting the sky to my liking, I added material below the two extensions I had made, which yielded the image below, which is now 2560x1728 (the image here is smaller). I can crop a 16:9 portion for my wallpaper folder, any size from 1920x1080 to 2560x1440.



It's notable that the extension of the village to the right in particular adds buildings that are not quite the same style as the original ones. Producing this image consumed 8 credits.

Second Session


I started a new session, uploaded the 1024x763 version and skipped cropping. Here I am about to add material at upper right. You can see that the uploaded image is smaller than a Generation Frame. I was hoping to produce a larger village, relative to the initial painting.

I wonder if I should have outpainted in smaller increments, such as having the Frame begin at the base of the large house. I did some Cancels this time. Buildings being added were often in a clashing style. Midway through the process I appended the words "finely detailed" to the prompt. It made little difference. Here we see the result after four corners were extended:


This image needs some work. As the next image shows, I did a lot of erasing while further extending. I ignored the added Moon; I plan to crop it out of a 16:9 image later.


Erasing while outpainting produced an image I like a bit better, shown next, but I like it less than the final image from the first session. 


This image (here it is reduced) is 2304 x 1792. I can crop out anything from 1920x1080 to 2304x1296 if I decide to use it as wallpaper. Producing it consumed 10 credits. Not as much bang for the buck as usual.

Third Session

I decided to try again with the 1024x763 version. Before outpainting I entered a different prompt: An old-time village in winter just after sunset with meltwater on snowy roads, in luminous cottage style

This time I extended downward first, to fill a 1024x1024 square, then I extended to the right and the left, with one Cancel, to get this:


This is 2240x1024. The added material is a bit cruder than the original painting, so I won't outpaint further. Rather, to produce a wallpaper image I'll increase the resolution (I use IrfanView, but Photoshop and Gimp are also good) to 2362x1080. I may use Unsharp Masking to boost the apparent sharpness a little. Then I can crop out a 1920x1080 piece, probably more to the left to include less of the reddish buildings. In the prompt I didn't refer to Geno Peoples, but used a more generic style note. I am not sure to what extent it helped. This image consumed 4 credits. I like it better than the one from Session 2.

Dall-E2 is an enjoyable collaborator, very useful but sometimes recalcitrant. It keeps things interesting.

Tuesday, August 29, 2023

Automated art, landscapes and beaches

 kw: experiments, artificial intelligence, simulated intelligence, art, generated art, images, prompts

I tried a few prompts to produce landscape paintings, even though both photos and scanned artwork, in large (HD and 4K sizes), abound in Google Images and Yahoo Images. I have downloaded many such images into folders that are used by a "screen saver" slideshow program when my computer is idle. Anyway, I thought it worth trying. I first tried Asian art, then more "ordinary" forest scenes, using the following prompts:

  1. A woodcut of a pagoda near a lake in mountainous country colored with pastels
  2. Ink and watercolor, highly detailed, Chinese scroll landscape with pine woods, steep hills, and a river
  3. a calming forest scene with wildflowers in a meadow, a stream, and a small pond, landscape painting
  4. A painting of a meadow in a hardwood forest with a stream and scattered wildflowers, mountains in the distant background


The prompts are clockwise from top left. The pagoda image is rather clumsy. There are two ways that block prints IRL are colorized. Some are printed from a single block with shading and then hand-colored later, and others are printed from multiple blocks with colored inks. This image looks like the latter case. That is to say, multi-block prints were used by Dall-E2 as the "universe" of images to respond to my prompt.

The scroll painting has interesting frame shifts in three places. Near center, what appears as a small waterfall at the base of a dark cliff morphs into a root wrapped on a broken-off tree trunk with growth on its top. At lower right, the scale seems to change, with tiny trees appearing to be far away when viewed in isolation. At upper left, a bit of detached landscape floats in the air. Dall-E2 does such "hallucinations" at times.

The two forested landscapes are both pleasing to me, and I often use one or the other as a background with using Zoom with a green screen. I've found that Dall-E2 "assumes" conifers are meant when "forest" is prompted, unless one specifies "hardwood", or perhaps "deciduous". These hardwoods appear to be birches. If I wanted oaks and maples I'd have had to say so.

I also like beaches, so I tried a couple, with these prompts:

  1. A sandy beach next to a rocky beach below cliffs, digital art
  2. A long stretch of sandy beach backed by sea cliffs, winding towards a distant horizon, digital art


The upper image reminds me of a portion of Pismo Beach, a place I stayed a couple of times several decades ago. The lower one came out like an aerial photo or drone shot. The shape of the bluffs is reminiscent more of a semi-tropical area such as New Zealand (sadly, known to me only through photos). I could have erased the sand bars at lower right and had Dall-E2 put in more water, but they do add a little interest. My favorite body-surfing beaches, near Huntington Beach south of Los Angeles and Sunset Beach to the north, have sand bars at low tide, which at high tide generate large, fast breakers that are a thrill to ride, but not popular with surfboarders.