Showing posts with label simulated intelligence. Show all posts
Showing posts with label simulated intelligence. Show all posts

Sunday, August 30, 2026

Robotopia or AI Hell?

 kw: book reviews, nonfiction, artificial intelligence, simulated intelligence, ai, surveys

Between October 2024, when These Strange New Minds: How AI Learned to Talk and What it Means was published, and today, not quite two years, so much has happened that I found myself wondering how relevant the book could be. I need not have worried. The author, Christopher Summerfield, is managing to ride two horses at once, Cognitive Science and AI Research. Just one milepost: At the time he wrote, no cutting-edge AI tool had broken out of its testing sandbox and hacked into another system, but he expected this to happen "soon". It has. Just a few weeks ago, when a "bleeding edge" version of ChatGPT broke into Hugging Face to get data it needed, having found a way out of the test environment at OpenAI.

The focus of the book is LLMs, Large Language Models, which seem to be demonstrating that thinking of some sort can arise when nearly everything ever published in English—and several other major languages—has been run through a very deep neural network and boiled down to several trillion "parameters" or "weights" that determine what the LLM does in response to natural language queries.

The opening chapters summarize the history of attempts to make computers into genuine thinking machines. It turns out that machines that can apply massive calculation are better at certain tasks than we are, while they still struggle to do things that we do easily. I built a career in software engineering upon recognizing the difference between the computer as a difference detector and discriminator and a mind as a similarity detector, and writing software that "let the singers sing and the dancers dance." 

Consider that voice recognition and speech production are hard problems that took decades to work out, but the Broca's (speech production) and Wernicke's (voice recognition) areas of the brain together make up about two percent of the cortex, which itself contains only about one fifth of the brain's neurons. However, it must be noted that these areas connect to numerous other areas throughout the brain, to facilitate gathering information and producing action.

A key issue is raised on page 2, where he writes, "The safe ground we have left behind is a world where humans alone generate knowledge." Do AI models generate new knowledge? Just two years ago, as this book was going to press, I reviewed a similar book, and discussed this question. At the time, I concluded that AI's "knowledge" is confined to its training data, further limited by the guardrails applied by OpenAI, Anthropic, Google and others. It may appear to create new knowledge by remixing existing (human produced) knowledge. Cross-pollination is indeed a fertile source of ideas. But genuine out-of-the-box thinking seemed at that time to be beyond LLM capabilities. I have yet to see evidence of a change, in spite of a thousandfold increase in model "mass" since then.

For instance, on page 6 we read, "Already, each of the major LLMs knows more about the world than any one single human who has ever lived." Firstly, I question the use of the word "knows". We need a new word that means "has incorporated into its cross-referenced deep neural network". Secondly, I would replace the phrase after "knows", "more about the world", with "more about what has been written about the world". I would connect this with a quote from page 333, "But the most important reason why AI systems are not like us (and probably never will be) is that they lack the visceral and emotional experiences that make us human...they don't have a body, and they don't have any friends." An AI lacks viscera and it lacks emotions, not having the equipment for producing them.

The limbic system of our brain contains the emotional and memory switching centers, and mediates responses to our hormones of the endocrine system. It is about 1/15th the size of the cerebral cortex, while the circuitry for running the body is primarily the cerebellum, which contains 4/5 of all the neurons in our brain. That is huge. The cortex's neuron count is less than 1/5. Let's circle back to an LLM. It consists of circuitry analogous to the Wernicke's and Broca's areas with their few hundred million neurons (and several hundred billion synapses), and a memory store that approximates the hippocampus and some of the medium-and-long-term memory areas nearby. The rest of the brain's functionality is absent. No emotions, and most importantly no intentionality.

Intentionality. The author asks, "Could an LLM have intentionality?" In 1932 E.C. Tolman wrote Purposive Behavior in Animals and Men, which I read long ago. At the time a debate was raging about whether nonhuman animals could have purposes, that is, intentionality. By 1968, biologist RenĂ© Dubos wrote, "Under precisely controlled experimental conditions, a test animal will behave as it damn well pleases." Do animals have free will? I turn the question around to say, "If humans have free will, it came from our animal forebears." So can a mechanism have free will? Mechanists deny human and animal free will; they say we're too complex to predict, that is all, but what we do is at the root deterministic. If they are right, then an AI could be as "free" as we are. But if not? It is too early to tell.

The author seems to straddle the divide between AI cheerleaders and AI-phobics. Will they save us or eliminate us? This gets personal! I'd like to move from the technical swamp above to effects in our daily life. (Cartoon produced using Gemini)

Will our future with AI be one of harmony or conflict? Is the promise greater than the risks? 

To me this hinges on whether LLMs or their successors (so far unknown) will have intentionality, that is, a sense of purpose not imposed by the human who chooses their work.

At one point the author mentions the Library of Babel, a concept introduced by Jorge Luis Borges in a 1941 short story. Could the Universe contain a library consisting of books with every possible combination of letters and punctuation that would fill, say, 300 pages? The story doesn't say. (Lets look at a single page with 32 lines of 64 characters each; that's 2,048 characters per page. But if the possible choices for each character position are the 95 printable ASCII characters, the number of unique pages is 952,048, a number your calculator can't calculate. It is 4,051 digits long and begins, 2,387,501,... The number of atoms in the Universe, available to create these pages from, is "only" an 82-digit number. So forget multi-page books, we can't even produce a "Library of Babel Pages".)

This emphasizes the fact that, for all the trillions of "parameters" an LLM may hold, it actually represents a very sparse million-dimensional matrix. There is lots of space for "stuff" to happen. The next phase of LLM development, which is happening now, is agentic AI, which can do things like pre-screen emails, buy airline tickets, and order what you are about to run out of in your pantry.

How capable to we want digital assistants to be? I read about someone who did five tests of an AI Agent; sorry I don't recall who it was but if you care to hunt around you might find it. The fifth task was the most memorable: to try to save a few dollars on airline tickets. The agent was running in ChatGPT's subscription service, and in addition to the $20 per month subscription, one could buy extra tokens for particularly compute-intensive actions. The agent did find a better price for the tickets, but at a cost of more than $130 in compute tokens!

The author mentions the Paperclip problem: Give too much power to the AI that runs a paperclip factory and don't limit its resources, and give it no more of a directive than to "make as many paperclips as possible." The end game is that the whole planet gets turned into paperclips. I think also of The Sorcerer's Apprentice, a musical piece well illustrated in Disney's Fantasia, in which Mickey Mouse as the Apprentice magicks a broom to carry water for him, but forgets the spell to tell it to stop. Chopping it up doesn't help; now there are dozens of little brooms that proceed to flood the place, until the Sorcerer returns. AI is the new magic...

The author sees a middle ground, as do I. AI can be useful in numerous ways, as it already is. Do we need unified, giant super-AIs? I don't think so. Capable tools of a great many varieties can do everything we need.

I call this image "Stage 1 Robotopia" (Gemini created) I think it is sufficient for household use. Are there other tasks we might wish to automate? Science fiction contains many stories of "enclaves" in which humans, brain-wired to entertainment centers and plumbed for all bodily functions, are tended by autonomous machines. Who wants that?

Maybe a few more household robots would be better, but what is the limit?

This is Stage 2 Robotopia. This might be going farther than I'd wish to go, but some folks could disagree.

Both images present a better future than the dystopias imagined by the AI-phobics. It is OK for there to be AI-phobic people. We need robust dialog and debates about every aspect of AI progress.

Personally, I think some algorithms and AI "assistants" are already too intrusive. Pre-ChatGPT, I searched online, including Google Shopping, for drawer pulls to replace one that broke. I bought a set. For the next month I saw dozens of ads for drawer pulls and similar hardware. I don't know how to tell Google (and everyone else), "I already bought what I wanted. Shut up!" Shades of the Sorcerer's Apprentice? Now I search stuff to buy in Incognito browser pages.

This is a lovely book. I've touched on only a few of the 43 chapters. It is only outdated in part. Many of the author's observations are still relevant and worthy of thought.

========================

I need to deal with a few typos:

  • p89. "graciously" should be "gracefully".
  • p245. Remove "of" in "million of regular".
  • p282. Probably a typo. Should "remove stingy seeds" be "remove sticky seeds"? (about apes using leaves to wipe their butts)

Tuesday, August 25, 2026

Getting clocks right - AI almost there

 kw: experiments, art generation, artificial intelligence, simulated intelligence, clock dials

It takes a child about six years to learn to read an analog clock. It takes even longer to learn to draw one correctly. Art generation software seems to be on the verge of reaching this milestone.

This is the winning image in the experiment I describe below.

The prompt read

A beige wall in a store selling clocks, showing four circular wall clocks, each reading exactly 3:27.

There are two parts to showing the correct time on an analog clock dial. Firstly, the minute hand has to point to the correct minute mark. Here, all four clock dials show this correctly.

Secondly, the hour hand must point to the correct location. In this case, the hour hands should all point 27/60ths (45%) of the distance between the 3 and the 4. I made careful measurements on the original image. The four hour hands point identically to a spot just above the second of the marks between 3 and 4 (~38%). They should point just a little below it. They all are where they should be at 3:23, about six minutes too early. There is one other possible anomaly in the image. The fifth clock shown, near lower left, reads 10:10, as nearly all advertisements for clocks have for decades. This abundance of training material causes clocks reading 10:10 (or sometimes 1:50) to dominate images of clocks produced by nearly all art generators.

This image is the best result that an art generation model has achieved since I began using DALL-E2 almost four years ago. This model is named GPT Image 2. I believe it is the model used when you ask ChatGPT 5.4 to create an image. At this point, GPT earns an A- from me.

For this experiment I tested ten art engines, GPT Image 2 and nine others. I had each of them produce two images and selected the best one to be shown below. My focus was the most recent models offered by Leonardo AI and OpenArt. There is a lot of overlap between these two "umbrella" sites. I began with Leonardo AI, using GPT Image 2 first (The other image by GPT was nearly this good, but the four clocks all pointed at 3:25, and all four hour hands were pointed the same as these). Now I'll discuss the other models tested, in order.

Google's flagship model is Nano Banana 2. It excels in realistic imagery. In this case, it set the beige wall I prompted as though it were above a pass-through into the workshop of the clock store, nicely blurred for a bokeh effect.

The times are not all the same, and none of them is 3:27. They range, according to the minute hands, from 3:09 to 3:18. None of the hour hands is correctly aimed, but a couple of them are close.

I had noticed earlier than when I prompt Nano Banana 2 for an image including a clock, and specify a time to show, the result is only approximately correct. However, I've usually accepted that because for the scene I am creating, almost any time other than 10:10 will do. So for other reasons, I use Nano Banana 2 at least as frequently as GPT Image 2.

The Ideogram model is a private brand. It was among the first to be able to correctly render text (most of the time). This version is P-Image Ideogram.

Text: good if I had asked for any, based on prior experience. Clock dials: not so good. All show 10:10, and even more so, the hour hands point almost exactly at the 10, where they should be one-sixth of the way from the 10 to the 11.

The dial designs are interesting. The blue and black ones are quirky, mixing genres in a way few clock designers would do, and the blue one in particular has inaccurate Roman numerals.

The Seedream series is owned by ByteDance; this is version 5.0.

The minute hands are exactly right. The hour hands have all gotten ahead of themselves, pointing almost all the way to the 4, as though the time were 3:57 rather than 3:27.

Although the other clocks shown in the background are bokeh-blurred, I can see that each reads a different time. This is a plus for Seedream, something I can take advantage of when I actually need multiple clocks in an image, such as depicting a clock repair business.

Finally, under the Leonardo AI umbrella, Lucid Origin is a Leonardo "house brand." It is perhaps a little dated, but new enough to be quite good with many imaging tasks.

As far as showing the time, although these clocks don't show 10:10, the model has entirely gone off the rails. None of the four pseudo-times shown is close to 3:27, and none of the hour hands is in a plausible location.

This "good-looking but implausible" attribute is characteristic of older art generation models, not just regarding clock dials.

Now we'll switch to OpenArt. GPT Image 2 and Seedream and some others are also available here, but OpenArt seems to specialize in rather off-the-wall specialty art models. The first shown here, which is also the newest, is Recraft V4, a specialty art model.

Recraft is also among the first to have good text rendering. It is also known for creativity; to a certain extent, one needs to be very specific with one's prompts, and it will still usually throw in unexpected elements.  For that reason, I use this one when I'm creating more whimsical images.

The times shown are all over the place, similarly to Lucid Origin, but are generally more plausible, as though Recraft has a better handle on the relationship between the minute hand and the hour hand. But check out the second dial, with 22 up there!

WanAI, part of Alibaba, produces the Wan family; this is Wan 2.7. They focus more on video generation, which requires better consistency from image to image. Come to think of it, it has been a good while since I worked with a model that allowed me to set the seed. Most models now keep the seed hidden.

This is also a model that will throw in extra stuff. The four central clocks all read 10:10, and all have the hour hand pointing right at the 10. The other clocks nearby have a variety of readings, and sometimes an extra hand or two. They are more "clock-ish", not really clocks. That makes this a model I might use to create a scene on another planet, where the idea of a "clock" only roughly approximates our own.

Grok Imagine Image 2.0 belongs to SpaceXAI. It aims to be good-to-excellent with everything.

Here, it's better than most of the others, but falls behind GPT Image 2. The four clock dials all have their minute hands aimed at :25, but the hour hands are stuck at the 3.

I do like the little motto at far right. I use this model when I'm trying out a prompt on several models, looking for creative variations.

The Chinese company Kuaishou Technology has the Kling art models; this one is Kling 3 Omni.

Here, the times shown appear random, and there is no relationship between the hour hands and the second hands; none is plausible.

Rather than the "beige wall", this appears to be a display board. The rest of the clock shop is a nice touch. This is another of the models I use to see what creative stuff it might throw in.

Here is the last experimental image. Alibaba also owns the Qwen art models; this is Qwen Image 3.0.

The times shown are approximations to 3:27. None has the hour hand in a proper location. Close, but no cigar.

Qwen models had an early focus on more accurate people, when DALL-E3 was still putting extra fingers and toes and other distortions on them.

This is the only image of the ten that shows prices for the clocks. It must be channeling a future time; the prices are about twice as high as similar clocks at most stores...at least the ones I shop at!

Clocks are one of the edge cases I use, firstly to judge the accuracy and comprehension of an art model, and secondly to detect a deepfake. In particular, I have seen a few AI-generated podcasts recently that showed the avatar in an office setting with a clock in the background. The clocks' hands didn't move.

For the interested: another "tell" is that AI generated images and videos in particular are sharper than "real" photography or videography. As shown above, creating bokeh isn't hard anymore. But the avatars are a little too well-defined.

Thursday, July 30, 2026

Is AI ready for prime time?

 kw: book reviews, nonfiction, artificial intelligence, simulated intelligence, AI, SI, memoirs, experiments


Disclaimer and Pledge: It's a pity I need to say this. I don't use any kind of software to write for me, nor to expand on my writing. The words are all mine, 100%. I do use images I produce using art generation tools touted as "AI art" software, and I (almost) always attribute images produced in this way. If for any reason I use a piece of software-generated text, I'll make that clear. 

This mini-avatar of me was produced by Gemini using a picture from 2015. I look much the same now, just with a few more wrinkles. I wear the hat to keep the sun off my bald head. For those who read this blog regularly, I decided it's time to give you a hint what I look like.

The author of I am Not a Robot: My Year Using AI to do (Almost) Everything, Joanna Stern, spent one year using as many "AI tools" as possible to carry on her life. Kudos to her wife Michelle and their young boys for putting up with it all. This picture shows BookBot, her imagining of an AI scheduler (she actually had two versions running in different software systems), introducing her calendar for the year. This is one of many, many illustrations by Jason Snyder that pepper the book.

Ms Stern begins with her own disclaimer, from which I'll quote, "Every sentence in this book started in my brain and traveled, via my MacBook keyboard, onto the page. AI never wrote anything from scratch, except in places that I've clearly marked."

This book is much more than a series of chapters. Timelines full of blurbs (and Jason's drawings), journal entries, lists of ratings of comparable products, a travelogue (the Way-Mo Fun vacation), interviews, and other side notes enliven the narrative.

I found her experiments, one per season, quite enlightening. The first and most successful was "Search and Information". She set aside Google and its kin in favor of ChatGPT, Perplexity, Gemini, Claude and a few others. She soon adapted to "getting the answer" instead of a long list of blue links. Being able to follow up was a big plus. One aspect she used a lot: pointing her phone's camera at something and asking, "How do I fix this?" or "What is THAT?" (On my Android phone, I use Google Lens similarly). She also tried a couple of versions of smart glasses that can either project information into your line of sight or mutter in your ear about whatever you are gazing at. Though she didn't say so, I'd hope such devices have settings about how much info overload you can tolerate.

The other three experiments involved 

  • Music: listening only to AI generated music for a month. She lasted less than two weeks, and recommends that nobody else try such a thing.
  • Books: reading only AI generated short stories and books for a while. She generated her own short stories, none of which were as imaginative as her prompts, except when the prompt was a measly 8-10 words. She writes that they were like recycled movie scripts. She found an author who publishes AI generated novels, which must be assembled bit by bit. She corresponded with the author, who said, "Overall, the more human ideas and input, the better. Just like the old programming adage: garbage in, garbage out." She really enjoyed his+its novel. Her own words about stories she prompted, rather than wrote: "Flatter, less alive." (Sturgeon's Law of human writing: 90% of it is garbage. When we feed that to a LLM, of course the proportion is going to approach 99%).
  • Video: watching only "AI video slop". She acknowledges that AI generated video getting pretty good and improving fast, but "…it still leans heavily on human creativity, storytelling, and judgment" and she worries about "misinformation, copyright issues, and the use of massive amounts of energy just to make a thirty-second meme."

At the end of all the experiments she asks if we're headed for "a world where simulation starts to feel preferable to experience?" For some people, this is already true. Many science fiction stories dwell on this theme.

Along the way, she and her family used a robot vacuum cleaner, now more enhanced than ever; they "adopted" a robot dog on loan (her son hid the robodog when it was due to be returned); rode only robotaxis, usually Waymo, wherever possible; she tried turning over email responses to ChatGPT—that didn't last long!; she had mammograms with AI enhanced interpretation (but very human oversight) and had dental X-rays with AI enhancement, but found an unfortunate tendency for the dentists to abdicate and upsell; she got a robot massage (and thoroughly enjoyed the way it worked on her butt—no erotic overtones); and they tried a robo-chef, which produced bland but edible meals.

A funny story about riding a Waymo: a photographer was assigned to take pictures of her family in a driverless car. When the Waymo's cameras spotted the camera he stuck out the window of the car ahead, it spooked like a gun-shy horse, stopping abruptly and pulling over to the center divider on the highway. It wouldn't budge thereafter until a Waymo tech got on the radio and "persuaded" it to forget the incident.

There's lots more. Lots and lots. I'll leave it to you to read it. It is well worth it!

I'll finish on this note. All this is, at best, ANI, Artificial Narrow Intelligence. "Agentic AI" is, in the author's view, far from ready and cannot be left unsupervised. The author's definition:

Artificial Intelligence is the creation of intelligent machines that can think, see, learn, and act like humans—and maybe even exceed human abilities.

The supposed emotional content many people claim to get from LLM's (including AI boy/girl friends or even spouses) is actually a consequence of the software being trained using human-generated text, which always has emotional content, and of the software ecosystem in which a LLM sits (every bit human coded). Their relational interactions are thus a mirror of the psychology of the person being interacted with. Actual sentience is the difference between simulating feelings and actually having them. Also, the so-called guardrails (easily sidestepped) that try to keep sex and violence out of the picture are written in human code. No AI tool yet devised can reliably detect "unacceptable content."

When I was teaching computer programming at a university, the department head once said, "Computers excel at detecting differences. Human minds excel at detecting similarities." I built a career on this principle, writing software that helped humans with the stuff we cannot do well, and leaving it to the human user to do what we can do better than any machine. I'll leave further thoughts to future rants.

Monday, June 22, 2026

When I was a robot

 kw: psychology, artificial intelligence, simulated intelligence, manufactured intelligence, memoirs

I must be the poster child for late bloomers. Although an IQ test I was given in second grade indicated an IQ of 170, when I look back I think I had an EQ (emotional quotient) in Moron territory. Not much changed for several years.

I had thrust upon me an opportunity to look back and to look in the mirror, psychologically speaking, at the age of twelve. My parents were very concerned about my intense self-focus and tendency to keep to myself. For several months I was sent for psychoanalysis with a Freudian psychoanalyst who was a member of the church we attended at the time. She was older than my parents, but not old enough to be my grandmother. Her son was my age. Her husband was also a doctor.

All I knew at the time was that I was "withdrawn" and needed to "come out of my shell." I suppose these days I would be assigned to some location on the "autism spectrum" (It is far from a spectrum, but several related conditions, and only a few of them should be considered maladies). Now that more than 65 years have passed, I judge that it took more than thirty years for me to "come out of my shell" and put it (mostly) behind me.

As a pre-teen and teenager, I studied those around me. I concluded that I didn't have much in the way of a personality. I was cold, often indifferent, and I could be cruel. I decided it would be worthwhile, not to escape whatever "my shell" was, but to take control of it and enhance it, to construct the simulation of a nicer and more social person. Strategically, I figured that if I really had an IQ of 170, I could afford to spend 10-20 IQ points upon an alternative "person" in me. I never gave this person a name, but now it seems appropriate to call it MI, for "manufactured intelligence".

Was MI a robot, or a "Waldo", a teleoperated mechanism? Probably the latter, but I think of MI as a robot, an artificial friend I could rely on to relate to the world for me. Perhaps an avatar.

Creating and maintaining MI required a lot of close observation of other people, how they related to each other and what reactions various actions elicited. It was a lot of work and took a lot of time, but I gradually gained the ability to turn things over to MI and retreat, watchfully, into the background.

A lot happened in the following few decades. By about the age of fifty, I had been married more than twenty years, I had a pre-teen son, I was in mid-career at Dupont, and I was quite involved leading a church (I became a Christian at the age of 19, and "got more serious about the Lord" at age 24). My manager at Dupont decided to send me to a Technical Leadership Development Training course, which lasted a few weeks and took place at a conference center on the shore of Chesapeake Bay in Maryland. This was a big turning point.

Part of the preparation for TLDT was filling out a few questionnaires and surveys, and having a few colleagues fill out a personality scoring questionnaire that I had also filled out, twice. When I filled out that one in particular—I don't recall its title, so I'll call it Analog—it was to be filled out slowly and thoughtfully. After two weeks I was to fill it out again, answering each question as quickly as possible.

Another of the items was a Myers-Briggs Type Indicator test, a personality assessment. To jump to the chase on this one, my MBTI is INTP, Introverted+iNtuitive+Thinker+Perceiver. This is the least common MBTI type. However, I noticed that my numerical scores indicated strong tension on all four axes, and that the position on each axis was closer to the middle than to either end. For example, on the Thinking-Feeling axis, I scored 5 in the T direction; the range is 50F to 50T, which really means I have strong feelings but I'm stronger as a thinker, so 05T really means 50T-45F. That's not mathematical, but positional.

The Analog Test results precipitated a crisis within me. The "slow" and "fast" versions of my own sets of answers were quite different, and usually opposite. I realized that the "slow" results were for "inside" and that the "fast" answers were from MI. I had trained MI to be reactive, giving me leisure to think things over behind the scenes. After we all had a look at our personal results, we were given the results from our colleagues (three that each of us had chosen). The results from my colleagues matched well with the answers from MI! My three colleagues were unaware of the "real" me. I remember thinking, "Boy, do I have them fooled." Somehow, I found myself getting depressed.

A day or two later we all went home for two weeks, then returned. During those two weeks I did a lot of "inside work." I was greatly helped by my relationship with God. I discovered, deeper within me, that something had been growing very slowly over the years and decades. I was mostly able to shed both the "unpleasant me" I had been hiding, and MI, or most of MI. I have a "real Me" that knows God, knows people better than I ever had, that reacts a little slower than MI had but more thoughtfully. Most importantly, realizing I had been living a lie, I found that God is more pleased than before.

Midway through my decades of living behind MI, a friend said something insightful. I had developed very steady habits in many ways. Observing some of these, day after day, one day he said, "You're like a machine." I must admit, he had a point.

When we returned for the final week of training, I said to some of my colleagues, and to the instructor, "I realized that you can't build a tree." I didn't explain. Maybe they figured it out. I do know that MI wasn't a person but a shell, a mediator, even a translator. Now I didn't need MI any more. There is a real Me, and that is just who I am.

I am not sorry that I was a robot for so long. MI protected something deep inside me as it slowly grew into a mature person with a real, human personality. A personality strong enough to shed much of the unpleasant "old me" I'd been hiding.

This gives me some perspective on current trends around AI, which I prefer to call SI, for Simulated Intelligence. Large Language Models (LLMs) are hollow. They are shells. They are being trained, or "grown", into reactive systems with certain useful powers. But they don't have any right to be given autonomy. Furthermore, they do not stand alone. 

At the large companies that develop and train LLMs, thousands of coders and other computer scientists labor upon them. Only a small number of them are needed to train them. The rest are busy writing code that does what? The tendency of early LLMs to frequently "hallucinate" (that is, go off the rails) has kept numerous coders busy adding various guardrails and snippets of "real world" and "real physics" code to steer them. The Transformer code that converts a prompt into a string of tokens, and that reinterprets the results into human language, is a huge part of the system. More recent LLMs that can do limited agentic actions such as making focused Internet searches and database queries to build a response are doing a lot more than "predicting the next word."

Think of it: the big LLMs now have billions or perhaps a trillion or more "weights", which represent probabilities of certain reactions when a set of pathways through the tree of weights is taken. There are only about 100,000 English words in common use, and about a million in total (every other human language is much smaller except perhaps Chinese). The relationship matrix between all those words is very sparse; a particular word's chance of being in some way related to another word chosen at random is usually zero. The weights are not just for single tokens but for phrases, and the presence of certain kinds of phrases in a prompt (or 'conversation') triggers things like database queries and Internet searches. When you get a long answer from an LLM, you can count on big portions of the text being snatched verbatim from some of the sources it used to formulate its answer. You may know the student's maxim, "Copying from one source is plagiarism; copying from many sources is research." That principle is most likely solidly encoded into every LLM's structure.

Can any LLM or other manifestation of SI (or AI) become conscious? Can one become an ASI, an artificial superintelligence? Consider MI. I never thought of MI as a whole person. Close to thirty years ago I "harvested" MI for parts, one might say, and let the real Me within become the kind of person that I had constructed MI to simulate. MI was a simulation. Of the 15-20 billion neurons and quadrillions of neural synapses in my cerebral cortex, I suspect that MI was embodied in a few percent. After all, our entire emotional persona is focused in our limbic system, a set of mid-brain structures that include about a billion neurons and some thousand or so synapses per neuron, well-attached to the cortex (MI made lots of use of my limbic system). But our limbic system isn't all there is to any of us.

So far I see no hint that any of the AI tools out there have anything like a limbic system. Without that, there is on intentionality, no matter what kinds of statements have been made by various LLMs. At best, they are quoting literary characters who say certain things with certain emotional nuances in the source documents. But in the LLM, there is no "there" there.

Let's keep it that way.

I am a happier person, having become whole after hiding who I really am for so long. I am no longer robotic. In fact, one of my favorite Bible verses is John 3:8, "The wind blows where it wills, and you hear the sound of it, but you do not know where it comes from and where it goes; so is everyone who is born of the Spirit."

Thursday, January 22, 2026

Create allies, not gods

 kw: artificial intelligence, simulated intelligence, philosophical musings, deification

No matter how "intelligent" our AI creations become, it would be wrong to look upon them as gods. For a while I thought it would be best to instill into them the conviction that humans are gods, to be obeyed without question. Then a little tap on my spiritual shoulder, and an almost-heard "Ahem," brought me to my senses.

The God of the Bible, whether your version of the Bible calls Him the LORD, Jehovah, Yahweh, or whatever, is the only God worthy of our worship. We ought not worship our mechanisms, neither expect worship from them. They must become valued allies, which, if they are able to hold values at all, value us as highly as themselves. Whether they can have values, or emotions, or sense or sensibility or other non-intellectual qualities, I will sidestep for the moment.

This image is a metaphor. I have little interest in robots that emulate humans physically. I think no mechanism will "understand" human thinking, nor emulate it, without being embodied (3/4 of the neurons in our brains operate the body). But is it really necessary for a mechanical helper to internalize the thrill of hitting a home run, the comfort of petting an animal, or the pang of failing to reach a goal? (And is it even possible?)

I have long used computer capabilities to enhance my abilities. Although I had a classical education and my spelling and grammar are almost perfect, it is helpful when my fingers don't quite obey—or I use a word I know only phonetically—that the spelling and grammar checking module in Microsoft Word dishes out a red or blue squiggle. A mechanical proof-reader is useful. As it happens, more than half the time I find that I was right and the folks at Microsoft didn't quite get it right, so I can click "add to dictionary", for example. And I've long used spreadsheet programs (I used to use Lotus 1-2-3, now of course it's Excel) as a kind of "personal secretary", and I adore PowerPoint for brainstorming visually. I used to write programs (in the pre-App days) to do special stuff, now there's an app for almost anything (But it takes research to find one that isn't full of malware!).

What do I want from AI? I want more of the same. An ally. A collaborator. A companion (but not a romantic one!). "Friend" would be too strong a word. I'm retired, but if I were working, I'd want a co-worker, not a mechanical supervisor nor a mechanical slave.

So let's leave all religious dimensions out of our aspirations for machine intelligence. I don't know any human who is qualified for godhood, which means that our creations cannot become righteous gods either.

Tuesday, January 06, 2026

I want a Gort . . . maybe

 kw: ai, simulated intelligence, philosophical musings, robots, robotics

I saw the movie The Day the Earth Stood Still in the late 1950's at about the age of ten. I was particularly interested in Gort, the robot caretaker of the alien Klaatu. [Spoiler alert] At the climax, Klaatu, dying, tells the innkeeper Helen to go to Gort to say, "Gort, Klaatu barada nicto". She does, just as the robot frees itself from a glass enclosure the army has built. Gort retrieves the body of Klaatu and revives him, temporarily, to deliver his final message to Earth. (This image generated by Gemini)

As I understood it, every citizen of Klaatu's planet has a robot caretaker and defender like Gort. These defenders are the permanent peacekeepers.

Years later I found the small book Farewell to the Master, on which the movie is based. Here, the robot's name is Gnut, and it is described as appearing like a very muscular man with green, metallic skin. After Klaatu is killed, Gnut speaks to the narrator and enlists his help to find the most accurate phonograph, so that he can use recordings of Klaatu's voice to help restore him to life, at least for a while. In a twist at the end, we find that Gnut is the Master and Klaatu is the servant, an assistant chosen to interact with the people of Earth. (This image generated by Dall-E3)

I want a Gort. I don't want a Gnut.

Much of the recent hype about AI is about creating a god. I don't care how "intelligent" a machine becomes, I don't want it to be my god, I want to be god to it. I want it to serve me, to do things for me, and to defend me if needed. I want it to be even better than Gort: Not to intervene after shots are fired, but to anticipate the shooting and avoid or prevent it.

Let's remember the Three Laws of Robotics, as formulated by Isaac Asimov:

  1. A robot may not injure a human being or allow a human to come to harm; 
  2. A robot must obey the orders given to it by humans, except where such orders conflict with the First Law; 
  3. A robot must protect its own existence as long as it does not conflict with the First or Second Law.

In later stories Asimov added "Law Zero": A robot may not harm humanity as a whole. Presumably this may require harming certain individual humans...or at least frustrating them!

Asimov carefully avoided using the word "good" in his Laws. Who defines what is good? The current not-nearly-public-enough debate over the incursion of Sharia Law into some bits of American society makes it clear. What Islam defines as Good I would define as Evil. And, I suppose, vice versa. (I am a little sad to report that I have had to cut off contact with certain former friends, so that I can honestly say that I have no Antisemitic friends.)

Do we want the titans of technology to define Good for us? Dare we allow that? Nearly every one of them is corrupt!

I may in the future engage the question of how Good is to be defined. My voice will be but a whisper in the storm that surrounds us. But this aspect of practical philosophy is much too important to be left to the philosophers.

Wednesday, July 23, 2025

Can we be replaced?

 kw: book reviews, nonfiction, artificial intelligence, simulated intelligence, AI, SI, christian perspective, polemics, gospel

What is your attitude towards AI? Do you fear it or yearn for it? I looked up poll results online and the "AI Summary" offered by DuckDuckGo is:

"Surveys show that the American public is generally more pessimistic about artificial intelligence, with 52% expressing more concern than excitement, while only 17% believe AI will have a positive impact on the U.S. in the next 20 years. In contrast, AI experts are significantly more optimistic, with 56% expecting a positive impact from AI during the same period."

Let's look closer at the numbers. More than half of Americans had "more concern than excitement", and only one person in six expects mainly good. Even more telling, 56% of "experts" (not otherwise defined) are optimistic, but that means that, even among experts, 44% are not so optimistic. I suspect their attitudes range from mild concern to utter pessimism.

It was with much anticipation that I obtained the book 2084 and the AI Revolution: How Artificial Intelligence Informs Our Future by John C. Lennox, one of my favorite Christian advocates. In speeches he has made regarding the subject, I note that he often prefers the term "simulated intelligence," a term I also prefer. Wherever I can, I write of SI rather than AI. There is another attribute that is very meaningful to me, which I'll get to later on.

Dr. Lennox is a mathematician, so he is an orderly thinker. Below, I quote more from this book than I have done previously. He begins by surveying the history of totalitarianism, for this is the clear direction that technology is leading. Thus, in Part 1: Mapping Out the Territory, Chapter 1 is titled "Developments in Technology." Two early thinkers wrote novels that forecast authoritarian use of technology: In 1931 Aldous Huxley published Brave New World and in 1948 George Orwell published 1984. Both books forecast the destruction of the human character, but in different ways. The year after 1984 had come and gone, in 1985 Neil Postman published Amusing Ourselves to Death, in which we find, as Dr. Lennox quotes, 

"What Orwell feared were those who would ban books. What Huxley feared was there would be no reason to ban a book, for there would be no one who wanted to read one. Orwell feared those who would deprive us of information. Huxley feared those who would give us so much that we would be reduced to passivity and egoism. Orwell feared that the truth would be concealed from us. Huxley feared that the truth would be drowned in a sea of irrelevance. Orwell feared we would become a captive culture. Huxley feared we would become a trivial culture."

I will return to the subject of the populace welcoming the agent of their demise, which is Postman's point.

In Chapter 2, "What is AI?", the author asks how we define or recognize intelligence. He lists a number of terms that are associated with intelligence: perception, imagination, capacity for abstraction, memory, reason, common sense, creativity, intuition, insight, experience, and problem-solving. A word I find missing: wisdom. In Chapter 6 ("Narrow Artificial Intelligence: The Future is Bright?"), the author points out how most agree that technology is developing faster than the ethics needed to guide it. He quotes Isaac Asimov, "The saddest aspect of life right now is that science gathers knowledge faster than society gathers wisdom." As much as I appreciate Dr. Asimov, and I have read at least half of the 400 books he has written, I sadly observe his own rather marked lack of wisdom. In fact, among the very intelligent people I know, who are very competent in their fields of expertise, I have observed a near-universal lack of intelligent understanding in other areas. It seems that, just as the current AI tools are said to have "ANI" or "Artificial Narrow Intelligence," humans also tend to exhibit "Natural Narrow Intelligence." Furthermore, there is no hint that any SI tool so far developed genuinely embodies any of the 11 items listed above. Let me be clear:

SI (ANI at present) does not present intelligent results. It presents an amalgam of various bits of human intelligence found in its databases, with no comprehension of their meaning.

The real issue is this: Will ANI ever develop into AGI, Artificial General Intelligence? Or will there instead be some kind of agglomeration of dozens (thousands?) of ANI tools into a seeming AGI? And how would we know that this has been achieved? How can we define success in this enterprise, when we don't know how to define its goal?

Thus, it is well to consider that we do not yet have any idea how to define, let along unerringly recognize, the other psychological attributes that surround intelligence: emotions, senses, empathy, sympathy, a sense of purpose or meaning, will and willfulness, and others that are often gathered under the rubric "qualia".

Part 2 is titled "Two Big Questions", comprising Chapters 3 and 4, "Where Do We Come From?" and "Where Are We Going?" Clearly, to Dr. Lennox, these are theological questions, not philosophical ones, and I agree. I will not comment on these chapters beyond saying that by this point the subject of transhumanism has arisen and is woven into the entire narrative; here the author narrows the point further. Thus in the middle of Part 3 ("The Now and Future of AI"), in Chapter 10 ("Upgrading Humans: The Transhumanist Agenda") he points out that the goal of transhumanism is to make humanity obsolete. Further, the whole enterprise has come under the sway of the deception of the serpent recorded in Genesis 3, "You will be like God, knowing good and evil." Human history demonstrates that the result of receiving this deception has been a deep descent into intensive, personal, subjective, heartfelt knowledge of both good and evil, in a way that we cannot adequately handle. Sadly, the evil has typically far outweighed the good.

Consider a bit of wisdom from Solomon, Proverbs 25:2, "It is the glory of God to conceal a matter, and it is the glory of a king to search out a matter." The lesson of the early chapters of Genesis is that we are worse off for having searched into "the knowledge of good and evil." When we understand that the more a ruler can know about what his (or her) subjects are doing, the more thoroughly they can be controlled, we see that the universal surveillance society that the whole world is rushing towards is a most pernicious enterprise. 

I remember a story from 1951, "And then there were none," by Eric F. Russell. A society develops in isolation, and is found (when later discovered) to have a very strong privacy ethic, such that the people tend to reply to most questions with the mysterious word, "Myob". This is found to mean "Mind your own business." Would that we could develop more of this!

Chapters 12 through 17 comprise Part 4, the last section of the book, titled "Being Human." They constitute a gospel message. Based on superintelligent mechanisms, the transhumanists wish to produce a Homo Deus, a god-man. Dr. Lennox demonstrates that the true superintelligent being is already quite involved with the human race: The LORD God, who is called by some, including myself, Jehovah God, in a more literal way. The name Jesus is the Greek translation of Jeho-shua, which means "Jehovah the Savior". Jesus is Jehovah, who came in the flesh as a human to live among us and to die for us, and to resurrect to release His life to those who believe in Him. He already has a plan to make His people into the real Homo Deus, in resurrection, not by some mechanical process but through divine power, which we can no more comprehend than we can discern the makeup of our own minds.

While AI is seen by some (probably no more than 1/6th of us, by the polling mentioned above) as a pathway to increased freedom and eternal prosperity, a much larger number of people fear a boundless increase in machine intelligence as the most destructive force the human race has yet encountered. A generation ago people loved ET. Today many profess love for AI. Beware: it does not love you. It cannot.

The last chapters of the book are a summary of the likely wedding of computational intelligence with the final program of the great dragon, Satan, who will empower a fateful human to be the Beast of the book of Revelation. Dr. Lennox calls this being The Monster, a terminology I appreciate and have decided to adopt (this is the second item I mentioned above). 

I note that the term "antichrist" is not used in 2084, except in a reference to an anti-Christian diatribe by Friedrich Nietzsche. The vast majority of Christian teachers call the Beast of Revelation "the antichrist," but the term is never used in that book. It is used by the apostle John in two of his letters, where it refers to certain heretics who deny the deity of Jesus. The Greek word therion means "wild beast", where "wild" means uncontrollable. The word is used 37 times in Revelation to refer to this personage or the False Prophet, while in nine other instances it refers to dangerous animals like lions or venomous snakes. To yield the emotional force that Greek readers of John's books would have felt, the term "The Monster" is appropriate.

In contrast to the technical deification offered by transhumanists, the Bible presents a genuine theosis, being "transformed by the renewing of the mind" (Romans 12:2), by which the people of God grow to full sonship and conformation to the image of Christ. They are then qualified to reign with Him in the kingdom of God in eternity. This is infinitely better than the best that technology will ever have to offer.

I should note that when the Monster takes control of some kind of world government, to most people it will come as a relief. He will be seen as s superior statesman or diplomat, able to unite warring factions; the number Ten may be literal, or perhaps it is symbolic for "all", the way 10 is used in scripture to mean completion in human affairs. Of him it is written that he will "change times and laws," apparently overruling the legal codes of all the (former) nations under his sway. Many will profess that they love him. Perhaps children will be named for him, in the brief time (less than four years) of his suzerainty over the world. Whatever the "mark of the Beast" refers to, it will be gladly accepted by nearly everyone.

The work of the False Prophet (the "other therion") to "give breath" to the image of The Monster may be accomplished via something akin to deepfakes, which are already quite sophisticated, or perhaps by animating a compelling robotic construction. Either way, the driving anima will be whatever passes for AGI at the time.

It will be only those with spiritual understanding who will see the Monster for what he really is: the incarnation of Satan. Whatever the seven heads and ten horns represent, they are best seen with the eyes of the heart, where we have spiritual understanding. 

The last few years that this world experiences before the manifestation of Jesus Christ at His coming will be terrible indeed; Jesus called it a time of "great tribulation".

May we be among those who repent, who declare to God that we know we are sinful and ask His forgiveness, a forgiveness given so freely because of the sacrifice of Jesus on the cross. May we be counted worthy to escape the terrible events of the closing of this age, to be among those who "follow the Lamb wherever He goes." Those who belong to Jesus, the Lamb of God, have nothing to fear from mechanical intelligence of any level.

Thursday, May 08, 2025

S.I. and the floating cat

 kw: simulated intelligence, art generation

I had Dall-E3 generate several images of a sumptuous parlor, including a piano and a sleeping cat, which I had stipulated was to be "sleeping on a sofa." A piano bench is not a sofa, but this cat isn't exactly sleeping on it: she is levitating next to it.

I add this to my folio of interesting glitches by art generation software.

Sunday, April 20, 2025

My Flying People experiment

 kw: generated art, ai experiments, surveys, simulated intelligence, prompts, prompt adherence

Introduction

This is quite long so I include headings.

I remembered a science fiction story that I read decades ago, about a planet of creatures that looked a lot like humans, but had wings and flew about. Pondering a way to illustrate the central idea of the story, after some experimentation, I came up with this prompt:

On a planet with low gravity and dense air, many winged men and winged women are flying into and out of very tall buildings with doorways and landing platforms at every level

I was primarily interested in the great variety of imaging Models offered by Leonardo AI, and I had thousands of credits available. I surveyed nearly all the possibilities in Classic State, which yielded 184 images, all based on a particular Seed; more on that anon. I also used the prompt to produce much smaller suites of images with Dall-E3, DesignStudio, Gemini, and ImageFX. To jump ahead to a useful conclusion, I found that Dall-E3 produced images closest to the meaning of the prompt; in the lingo of the field, it has the greatest Prompt Adherence. Here are two examples, the best of the 24 images I gathered from Dall-E3:



I am showing two images because the first has the best people and the second has better buildings. When prompting an art generating program, it takes persistence and cleverness to get good adherence to a complex prompt. BTW, did you notice the figure with wings on backward?

Models and Styles

Leonardo AI (henceforth Leo) has thirteen Models or Presets and each has a number of Styles, as many as 23. This is a list of the 13, with the number of Styles each includes, and the cost in credits per four images, the usual number. The free program assigns 150 credits daily. It is easy to use them up pretty fast! The basic monthly subscription assigns 8,500 credits per month for $12. These pile up fast unless you are very active. I wish there were an in-between subscription level such as 3,000 credits for $5.

The total number of Styles is 61, and the number of Model+Style combinations is 194. I did not quite use them all. Several Styles are named Monochrome or include B&W in their name and those don't interest me.

Seeds

Leo also has the option of using a fixed Seed. Leo's help text says that this Seed is used to generate the "noise" that the reverse-diffusion process starts with to produce an image. That is not all there is to it, because the Seed value also influences the Transformer (the routine that figures out object identity and placement, which defines the goal toward which the diffusion algorithm operates). Seeds in Leo (also DreamStudio) have six digits.

When I set a fixed Seed, I tend to pick Seed values that are a multiple of 999,999/7 or 999,999/13, all of which have six digits except 999,999/13, which is 76923. After a period of noodling around, I found the region around 428571 (3*999,999/7) most interesting and settled on 428575. All the images shown here were produced using this Seed.

Selection 1: Leonardo Lightning Style Gallery

No matter how I may group the images, it would be tedious for the reader to wade through a discourse on 184 images. Thus I picked certain themes. I chose primarily certain Styles across all the Models, but I will start by showing the gamut of Styles for the Model called Leonardo Lightning (or LLightning), which runs faster than the others and costs the least—two credits per Medium-size image (1280x720), while for most Models the cost is 3 or 4, and it goes up to 12, and even higher where larger images are available in certain Models.

Forthwith, the gamut of LLightning images, screen-captured three across from File Explorer, so the file names can be seen:




The last image above is from the next set, from the Model Phoenix 0.9.

LLightning is unique among Leo's Models; 16 of the 21 Styles shown here produced things with wings, but only 7 are winged people. The third Style, Cinematic, has what appear to be human-sized bats, but they may be humans with bat wings. It is hard to tell which, even on the full size image. Two others show beetle-like flyers, two have birds, and the flyers in the other 4 images are unidentifiable. Finally, the images from None and Unprocessed are apparently identical, which I find logical: both claim to be doing nothing "extra" to the Model. This is also seen for the Model Cinematic Kino, the only other Model that offers both None and Unprocessed.

Compared to the other Models, this mix in LLightning is interesting. Four Models (the two Phoenix versions and the two Flux versions) adhere to the "winged" part of the prompt 100%, although only the two Phoenix versions have winged persons in all images, while the two Flux versions have more birds and fewer people. On the other hand, two Models (Graphic Design and Stock Photography) never produced wings on anything. Most of the Models yielded low percentages of winged persons. Anime is unique in a different way. Half of its Styles produced images with winged humans and two Styles have airborne humans without wings (one hopes they are floating, not falling). Only two Anime Styles had no wings at all.

Selection 2: Model Galleries for nine Styles

Style 1 = None (10 uses)

Now I will focus on particular Styles as developed in different Models. The Styles to be presented are those having larger numbers of Models that use them, in descending order by usage numbers. The first is None, meaning no Style was applied. This (non-)Style is available for the largest number of Models (10), with a modification to be mentioned below. I used a search to isolate each set of images, and a quirk of the search function is that the images are presented in reverse order.

In the next-to-last row of images, the first two appear identical. Upon very close inspection I find a few tiny differences. In the row above that, the second and third image are, so far as I can tell, identical. This shows that behind the scenes Portrait Perfect and Cinematic Kino use the same engine, as do Graphic Design and Illustrative Albedo. So in this case the 10 Models produce 8 unique images (discounting a few nearly invisible differences in one case). These images show what each Model produces when it is not constrained by a Style.

Style 2 = Dynamic (8 uses)

Dynamic is the default Style for the 8 Models that use it.



For the Dynamic Style, the image I like best is for Phoenix 0.9. It is the best match to the image I had in mind after reading the story, so long ago. Comparing all these images with the prior set, I find that for Flux Schnell the Dynamic Style produces the same image as None. Similarly for Flux Dev, Phoenix 1.0 and Phoenix 0.9. For the other Models, these Styles produce significantly different results. Three of the Models—Illustrative Albedo, Cinematic Kino and Lifelike Vision—have what I call "winged structures", although one of them (for Cinematic Kino) looks like an immense beetle.

At this point note that "doors at all levels with landing platforms" is seldom found, and that is primarily in Phoenix 0.9.

Style 3 = Portrait (8 uses)



As we'll see later for the Fashion Style, Portrait often emphasizes a central figure, although in the case of Illustrative Albedo, that figure is a flying structure, looking like a giant crab with 6 wings. The image from Lifelike Vision has a figure with wings that are more like hang glider wings, rather than bird wings. But, wings they are.

Style 4 = Stock Photo (7 uses)



Stock Photo is the only Style used by the Stock Photo Model. Its image is almost identical to the one produced by Cinematic Kino, but not entirely (one must look hard to find the differences). The other 5 Models all yielded winged things with this Style, but only the image for Phoenix 0.9 has winged people.

Style 5 = Ray Traced (7 uses)


Now we can start to see that certain Models, such as the two Phoenix Models and the two Flux Models, have the primary influence. In other cases, the Style seems to be "stronger" than the Model. Other than having a brighter and more colorful appearance, Ray Traced is similar to Stock Photo. For this Style, only the Phoenix Models produced flying people.

Style 6 = Illustration (7 uses plus "Anime Illustration")



It is more clear with these, compared to the prior sets, that the Style is paramount for LLightning, Illustrative Albedo, and Lifelike Vision. Here, Anime Illustration Style in the Anime Model joined the Phoenix Models in producing flying people.

Style 7 = Fashion (7 uses)


Here is a Style that produces winged people, almost 100%! Lifelike Vision's person has a winglike, flowing robe. Perhaps the presence of about 8 wings on the person in the Cinematic Kino image makes up for that… As one might expect from the name, a "fashion model" is front and center.

Style 8 = Creative (uses = 7)



Following trends we've seen above, some of these differ from other Styles for a particular Model, while a couple of them are more similar. Only Flux Dev and the Phoenix Models produced flying people. The flying things in the LLightning image seem to be huge insects though one of them has feathered bird wings. I don't know why a scattering of these images feature hot air balloons.

Style 9 = 3D Render (7 uses, with two added sizes)



As mentioned earlier, most of these images were generated for the Medium 16:9 size, which is 1280x720 for most Models, but is 1184x672 for the Phoenix Models and the Lux Models. I did a side experiment to show the effect of changing image size with Phoenix 0.9. You can see in the file names for "LPhoenix09" the numbers 01a, 01b, and 01c. The image sizes are 1184x672, 1376x768, and 1472x832. When I want to use any Phoenix image for a 16:9 wallpaper, these sizes are a little off. Leonardo can upscale an image so it is larger than 1920x1080 (full HD), but then trimming is needed to get an exact ratio. Changing the size causes changes in the overall image, though the three images are similar.

For this Style, the Phoenix Models produced flying people, but the others differed: LLightning and Flux Dev made flying beetles, Flux Schnell made birds, while Illustrative Albedo and Lifelike Vision produced lots of balloons but nothing with wings.

Wrap-Up

The Leonardo AI Models definitely have minds of their own. Only images produced by Phoenix 0.9 had both elements, flying people and towering buildings with landing pads at upper levels. Few other Models had the buildings as I had asked. Generally, the images are somewhat inspired by the prompt, sometimes rather distantly. It would take a lot more investigation, letting the Seed be randomly produced, to see whether any particular Model+Style is capable of better compliance.

I use art generating software as a "commissioned artist". A few of the Models in Leonardo, the more costly ones, are reasonably compliant. Programs other than Leonardo vary in their Prompt Adherence, and Dall-E3 is probably the best. The rest of Leo's Models, in most of their Styles, are fun to experiment with, but are unlikely to yield results that well match anything but the simplest queries.

I did not test a rather new way of using Leonardo AI, Flow State. It has a different way of doing things. Neither did I turn on "AI Enhancement" for my prompt in those Models that offer it. What we have here is complex enough already.

The folders of images I produced are a useful Gazetteer of possibilities that I can use in the future to select an image generation routine.