Showing posts with label computer science. Show all posts
Showing posts with label computer science. Show all posts

Wednesday, September 04, 2024

If it is artificial, is it intelligence?

 kw: book reviews, nonfiction, computer science, artificial intelligence, simulated intelligence, surveys

Before "hacker" meant "computer-using criminal", it meant "enthusiast". Many early hackers spent their time obsessively doing one of two things: either writing a new operating system, or trying to write software to do "human stuff". I have been hearing about "AI", "artificial intelligence" since the term was coined by Claude Shannon when I was nine years old. Two years later the third book in the Danny Dunn series was Danny Dunn and the Homework Machine (by Abrashkin and Williams). It featured a desk-sized computer with a novel design, created by a family friend, Professor Bullfinch. A decade later (1968) I had the chance to learn FORTRAN, which kicked off a lifelong hobby-turned-profession. The computer I learned on was the first desk-sized "minicomputer", the IBM 1130.

ENIAC and other early "elephants" were called "electronic brains" almost from the beginning. I learned how CPU's (central processing units) worked, and even though the operation of biological brains was not so well known yet, it was clear to me that computers worked in an utterly different way.

Fast-forward a few decades. Some 20 years ago a "last page" article in Scientific American described the newest supercomputer, claiming that it was equivalent to a human brain, in memory size, component count, and computing speed. Where it did not match the brain was the amount of power it needed: three million watts. Our brains use about 20 watts. It soon became evident that the metric of brain complexity is not the number of neurons, but the number of synapses, plus other connections between the neurons and the glia and other "support cells". In this, that supercomputer was woefully lacking. This is still true.

Tell me, does this illustration show one being or two?

My generation and all those following have been influenced by I, Robot, and by Forbidden Planet, and by other popular depictions of brainy machines. We think of a "robot" as a mechanical man, a humanoid mechanism that is self-contained.

In order to behave and respond the way one of the robots in I, Robot does, a humanoid mechanism would need an intimate connection with a room full of equipment like the supercomputer in the background of the image (that part of the image is real). When the Watson supercomputer played Jeopardy (and won) a few years ago, what the audience didn't see was the roomful of equipment offstage. And that was just what was running the trained "model" and its databases and the voice interface. The equipment used for training Watson was much larger, and kept a couple of hundred computer scientists, linguists, and other experts occupied for many months.

Assuming Moore's Law continues to double circuit complexity (per cubic cm) each two years, it will take sixteen doublings, or 32 years, to get the supercomputer shown into a unit that fits inside a robot of the size shown. And power requirements will have to drop from millions of watts to 100 watts or less. And this is still not a machine that has the brain power of a human. We don't know what that would take.

All this is to introduce a fascinating book, The Mind's Mirror: Risk and Reward in the Age of AI, by Daniela Rus and Gregory Mone. While Professor Rus is a strong proponent of AI and of its further development, she is more clear-headed than the authors of most books on the subject. In particular, she sees more clearly than most the risks, the dangers, of faddish over-promotion and of rushing blindly into an "AI Future".

At the outset, in Chapter 1, "Speed", she clearly emphasizes that products such as ChatGPT, DALL-E, and Gemini are tools, and particularly that their "expertise" is confined to the material that was used to train them. She writes that it might have been possible to get one of the LLM's (large language models) to write a chapter of the book, but it "would not really represent my ideas. It would be a carefully selected string of text modeled on trillions of words spread across the web. My goal is to share my knowledge, expertise, passion, and fears with regard to AI, not the global average of all such ideas. And if I want to do that, I cannot rely on an AI chat assistant." (p. 11) In a number of places she calls AI software "intelligent tools."

She continues the theme, writing of knowledge, insight, and creativity (Chapters 2 – 4), saying at one point, "They are masters of cliché." (p. 60) Critical analysis skills that we used to learn were based on following the progression from Data to Information to Knowledge, and then to Insight and Wisdom (does any school still teach this?) All of these together add up to comprehension. Does anyone have the slightest idea how to bring about Artificial Comprehension?

None of the software tools has shown the slightest ability to step outside the bounds of their training data. If ChatGPT "hallucinates", it is not rendering new knowledge, but remixing biased or deceptive content from its poorly curated training set, perhaps with a dollop of truthful "old news" in the mix. This illustration of a LinkedIn post I wrote last year shows the point.

The colors are significant:

  • Green = correct or true
  • Lighter orange = incomplete or outdated
  • Darker orange and red = false to malicious, even evil
  • Blue = AI training data, partly "good", partly "poor", partly "evil"—we hope not too evil
The three lavender blobs at right are varieties of human experience, including someone creating new "good stuff", poking out to the right and increasing the store of published knowledge. I kept those blobs far from the training data on purpose. Training is typically done with "old hat" material.

This book has a rare admission that "it's essential to remember that the nature and function of AI parameters and human synapses are vastly different." (p. 106) We don't know all that well the amount of processing going on within a neuron, nor even if a synapse is more than just a signal-passing "gate" or is something more capable. 

And though the matter of embodiment is touched upon, I was disappointed that there wasn't more on this. Perhaps you've heard that the human brain has "100 billion neurons". The actual number is 85-90 billion, and 80% of them are in the cerebellum, the "little brain" at the back, above the brain stem. We have a little inkling that the sorts of processing that cerebellar neurons perform are different from those in the cerebral neurons (the famous "gray matter"). Clearly, when 80% of the neurons make up only 10% of the brain's total volume, these neurons are smaller. The cerebellum "runs the body", and handles the traffic between the body and the cerebral cortex, the "upper brain", where thinking (most likely) occurs. Embodiment is clearly extremely important for a brain to function properly. It's being glossed over by most workers in AI.

A special chapter between 11 and 12 is "A Business Interlude: The AI Implementation Playbook". An entrepreneur or business leader who wants either to initiate a venture that strongly relies on these tools, or who wants to add them to the company bag of tricks would do well to extract these 19 pages from the book and dig into them. They include the key steps to take and the crucial questions to ask (including "Will AI be cost-effective compared to what I am doing now?"). A key component of any team tasked with making or transitioning a business to use AI is "bilinguals", people who are well versed in the business and also in computer science and AI in particular. This is analogous to a key period in my career (not with AI, though): Because I had studied all the sciences in college, and had a few degrees, and I also was a highly competent coder, I was a valued "bilingual", getting scientific software to work well for the scientists at the research facility where I worked. Bottom line: You need the right people to make appropriate use of AI tools in your company.

The book includes a major section on risks and the defenses we need. Whether some future AI system will "take over" or subjugate us is a far-off threat. It is not to be ignored, but the front-burner issues are what humans will do with AI tools that we need to be wary of. Something my mother said comes back to me. I was about to go into a new area of the desert for a solo hike. She said she was worried for my safety. I said, "Oh, I know how to avoid rattlesnakes." She replied, "I am worried about rattle-people!"

Let's keep that in mind. In my experience, rattle-people are a big risk whenever any new tool is created. What's one of the biggest uses of generative art-AI? Coupling it with Photoshop to produce deep fake pictures. Deep fake movies are a bit more difficult and costly just now, but just wait…and not very long! Soon, it will take a powerful AI tool to detect deep fake pix and vids, and how will we know that the AI detective tool is reliable and truthful?

A proverb from my coding days, "If we built houses the way most software is written, the next woodpecker to come along could destroy civilization." Most of us old-timers know a dozen ways to hack into a system, but the easiest is "social engineering," finding someone to trick into revealing login credentials or other information to help the hacker get into the system. Now social engineers are using AI tools to help them write convincing scripts to use to fool people, whether through scam phone calls, phishing emails or smishing SMS (or WhatsApp or Line or FB, etc.) messages.

[You can take this to the voting booth: Effective right now, any TV or radio political ad, particularly the attack ads, will have AI-generated content. If you want to know a candidate, go to the candidate's web site and look for actual policy statements (NOT promises!).]

A final matter I wish Professor Rus had included: Human decision making requires emotion. Persons who have suffered the kind of brain damage that "disconnects" their emotions become unable to make a decision. Somehow, we have to like something in order to choose it. Where "liking" comes from, we haven't a clue. But it is essential!

There is much more I could go into, but this is enough, I hope, to whet your appetite to get the book and read it.

A final word, that I didn't want to bring into the earlier discussions. I don't like the term Artificial Intelligence, nor AI. I much prefer Simulated Intelligence, abbreviated SI. It is unfortunate that, in the world of science, SI refers to System Internationale, the system of units such as meter, kilogram and second, used to define quantities in mathematical physics. Perhaps someone who reads this can come up with another moniker that makes it clear that machine intelligence isn't really intelligent yet.

Monday, April 24, 2023

BARD figures it out

 kw: analysis, artificial intelligence, experimental linguistics, mathematics, computer science

Continuing my exploration of the capabilities of BARD, I asked it about some computational and mathematical matters, starting with one that was relevant to recent work I have been doing. I'll comment along with each interaction this time.

P07: Two angles are related by the function A = arctan(sec B * tan B). Given A, what is B?

Bard: The given function is not invertible, so there is no unique value of B that corresponds to a given value of A. This is because the tangent function is periodic, with a period of 180°. This means that for any value of A, there are infinitely many values of B that will satisfy the given function.

For example, if A=45∘, then B could be any value in the range (−45∘, 45∘). Or, if A=90°, then B could be any value in the range (−90°, 90°).

In order to find a specific value of B that corresponds to a given value of A, you would need to specify additional information about the problem. For example, you might know that B is an acute angle, or that B is a right angle.

P07: In the above problem, both A and B are in the range 0° to 90°.

Bard: In that case, there are two possible values of B for any given value of A. For example, if A=45°, then the two possible values of B are 45° and 135°. This is because the tangent function is positive in the first and third quadrants, so the two possible values of B are the two angles in the first and third quadrants that have a tangent of 1.

In general, the two possible values of B can be found by solving the equation

[recast by P07]                  tan B = ± tan A / sec B

For example, if A=45°, then this equation becomes

[recast by P07]                  tan B = ±tan 45° / sec B = ±1 / √2

Solving this equation, we find that the two possible values of B are 45° and 135°.

The original statement is correct, though not only for the reason BARD presented. This is an example of a Trapdoor Function. If angle A is known, it is straightforward to calculate angle B. But if only B is known, finding A requires an iteration, such as successive applications of Newton's Approximation.

BARD treated this like a Diophantine problem, and ignored the restriction of B to (0°, 90°). Further, the answer is incorrect. Tan 45° = 1, sec 45° ≈ 1.414, so 1 is being equated to √2.

P07: Are you able to chart a function?

Bard: Yes, I can chart a function. I can use the following steps to chart a function:

1.       Identify the domain and range of the function.

2.       Choose a few values of x and calculate the corresponding values of y.

3.       Plot the points (x,y) on a graph.

4.       Connect the points with a smooth curve.

For example, let's chart the function y=x2. The domain of this function is all real numbers, and the range is all non-negative real numbers. We can choose a few values of x, such as x=0,1,2,3, and calculate the corresponding values of y, such as y=0,1,4,9. We can then plot the points (0,0), (1,1), (2,4), (3,9) on a graph and connect them with a smooth curve. The resulting graph is a parabola that opens upwards.

P07: Can you draw the chart?

Bard: I can't draw a chart for you here, but I can give you a description of the chart of y=x2.  The chart is a parabola that opens upwards. The vertex of the parabola is at (0,0). The x-axis is the line of symmetry of the parabola. The parabola approaches the x-axis asymptotically as x approaches ±∞.

This is just as I expected; BARD is linguistic in nature, and has no graphical capabilities...yet!

P07: What is the integral of Sqrt(sec(x)*tan(x))?

Bard: 



The line above needed to be a picture.

To solve this, we can use the following steps:

  1. Let u = sec x. Then du = sec x tan x dx.
  2. Substituting gives us:

Code snippet

∫√sec(x)tan(x)dx = ∫√u du = 2√u + C = 2√sec(x) + C

where C is an arbitrary constant of integration.

P07: Can a computer store the exact value of an irrational number?

Bard: No, a computer cannot store the exact value of an irrational number. This is because computers store numbers in binary, which is a finite representation of numbers. Irrational numbers are infinite, so they cannot be represented exactly in binary.

For example, the number π is irrational. It is a non-repeating, non-terminating decimal number. This means that there is no finite sequence of digits that can represent π exactly. If we try to store π in binary, we will only be able to store an approximation of it.

The same is true for all other irrational numbers. They cannot be stored exactly in binary, so they can only be approximated.

However, this does not mean that computers cannot use irrational numbers. They can still perform calculations with irrational numbers, as long as they are aware that the numbers are being approximated.

The only quibble I have with this answer is the word "they" in the last sentence. Computers aren't aware of anything. It should read, "…as long as those using the program are aware…"

P07: What is Green’s Theorem?

Bard: Green's theorem is a theorem in vector calculus that relates a line integral around a simple closed curve C to a double integral over the plane region D bounded by C. It is the two-dimensional special case of Stokes' theorem.

In mathematical terms, Green's theorem states that

Code snippet

∫CF⋅ds=∬D(curl F)⋅k dA

Use code with caution. Learn more [link to https://bard.google.com/faq#coding]

where

  •         F is a vector field defined on a two-dimensional domain D
  •         C is a simple closed curve in D
  •         ds is the line element along C
  •         dA is the area element in D
  •         curl F is the curl of F, which is a vector field that measures the circulation of F around a point
  •         k is the unit vector in the z-direction 

Green's theorem can be used to solve a variety of problems in physics, engineering, and mathematics. For example, it can be used to 

  •         Calculate the flow of a fluid through a region
  •          Determine the area of a plane figure
  •          Find the centroid of a plane figure
  •          Solve boundary value problems

 Green's theorem is a powerful tool that can be used to simplify and solve a wide variety of problems in vector calculus.

BARD then referred to the Wikipedia article on Green’s Theorem, and suggested three Google searches on the topic. I then asked about one term used in the description.

P07: What is the curl of a function?

Bard: In vector calculus, the curl is a vector operator that describes the infinitesimal circulation of a vector field in three-dimensional Euclidean space. The curl at a point in the field is represented by a vector whose length and direction denote the magnitude and axis of the maximum circulation. The curl of a field is formally defined as the circulation density at each point of the field.

In other words, the curl of a vector field is a vector that measures the rotation of the field around a point. The curl is zero at points where the field does not rotate, and it is nonzero at points where the field does rotate. The direction of the curl vector indicates the direction of the axis of rotation, and the magnitude of the curl vector indicates the magnitude of the rotation.

The curl of a vector field can be calculated using the cross product. The cross product is a mathematical operation that takes two vectors and produces a third vector that is perpendicular to both of the original vectors. The direction of the cross product vector is determined by the right-hand rule.

The curl of a vector field can also be calculated using the curl operator. The curl operator is a mathematical symbol that represents the curl of a vector field. The curl operator is written as a counterclockwise cross product, and it is applied to the vector field.

The curl of a vector field has many applications in physics and engineering. For example, the curl of the velocity field of a fluid is used to calculate the vorticity of the fluid. The vorticity of a fluid is a measure of the rotation of the fluid. The curl of the electric field is used to calculate the magnetic field. The magnetic field is a vector field that describes the direction and magnitude of the magnetic force.

The curl is a powerful tool for analyzing the behavior of vector fields. It can be used to calculate the rotation of a vector field, the vorticity of a fluid, and the magnetic field. The curl is an important tool in physics, engineering, and other fields.

BARD referenced a Wikipedia article on the subject. In an alternate draft BARD added:

The curl of a vector field F is denoted by ∇ × F. In three dimensions, the curl can be computed using the following formula:

Where F1, F2, and F3 are the components of function F.

Not being a vector calculus expert, I had to ask for the definition of the inverted delta.

The symbol ∇ is called the nabla symbol. It is a triangular symbol resembling an inverted Greek delta: \nabla or ∇. The name comes, by reason of the symbol's shape, from the Hellenistic Greek word νάβλα for a Phoenician harp, and was suggested by the encyclopedist William Robertson Smith to Peter Guthrie Tait in correspondence.

In mathematics, the nabla symbol is used to represent the gradient operator in calculus. The gradient is a vector that points in the direction of the greatest rate of change of a function, and its magnitude is equal to the magnitude of the rate of change.

The nabla symbol can also be used to represent other vector operators, such as the divergence and the curl. The divergence is a vector operator that measures the outward flux of a vector field from a surface, and the curl is a vector operator that measures the circulation of a vector field around a point.

The nabla symbol is a powerful tool that can be used to represent a variety of vector operators in mathematics. It is an important symbol in vector calculus and has many applications in physics, engineering, and other fields.

Sources:

  1. brainly.in/question/12687051 
  2. en.wikipedia.org/wiki/Nabla_symbol 
  3. mathoverflow.net/questions/289607/mathematical-objects-whose-name-is-a-single-letter  

BARD did what I would have done, look these things up in Wikipedia and math sources such as those it listed following its response. Green's Theorem was the bugaboo of a Senior-level course in Physics I took many, many years ago.

Here we see BARD doing less linguistic generation and more copying of sources.




Friday, July 13, 2012

Maybe worse than impossible

kw: programming, computer science

A couple of decades ago I was working on a program to simulate the way oil fills up a "reservoir", really an impermeable surface with some shape that could trap oil as it trickled upward. A colleague and I tried out scheme after scheme, mainly based on how hard each one was to write into computer code. Of course the easiest one to write elicited my colleagues comment, "Boy, that is pretty bogus." He meant that whether it worked or not, it was a very inefficient way to proceed. We eventually found a pretty efficient method. Other projects throughout the years have been, sometimes a search for efficiency, and sometimes a frantic scramble to avoid too much bogosity.

Definition: Bogosity in computer programming is the inverse of efficiency. Where "bogus" usually means "fake", to a programmer it means a bad way to do something, one that will make the computer take too long. But we need to see just how bad "bad" can be. The standard programmer's example is sorting, so I'll use it.

Sorting data has received huge amounts of attention from thousands of bright people because it is needed so frequently, and is very slow unless clever methods are employed. Sorting bogosity is easy. The (almost) easiest method is one you can do by hand, and it matches the way people typically operate when sorting a small number of things, such as lining up a dozen rocks from smallest to largest. This easy method is called the Bubble Sort:
  • Line up the rocks and look through them for the largest. Put it at the end of the line.
  • Look for the largest among those that remain. Put it next to the other one.
  • Repeat.
We can do this rather quickly because we can compare several items at once by eye. But in a computer, it must be done by comparing them pairwise. So a computer would have to use a few extra steps:
  • Load numbers representing the "size" of the rocks (such as weight. That means weigh them first).
  • Compare the first two numbers. If the first is bigger than the second, swap them.
  • Compare the next two numbers, and so forth, swapping as needed, until you get to the end of the list.
  • At that point, the last storage location contains the largest number.
  • Start over, but stop one short of the end. Now you have two numbers in order.
  • Repeat until you have made 11 passes through an ever-shortening list of numbers.
For 12 rocks, the program made 11 comparisons, then 10, and so forth, for a total of 66. That's OK for a dozen items, but what about sorting a deck of 52 cards? Once you determine which order the suits go in (whether to sort by suits or by numbers within suits), you have to make 1,326 comparisons. That is a lot for a human, though it is pretty quick for a computer. But computer programs need to do a good job even if sorting tens of thousands of items. For 10,000 items, the bubble sort takes almost 50 million comparisons.

A different kind of sort works better. I'll describe one variation of the Shell Sort. It takes fewer comparisons, but uses some extra space. First, for 12 rocks:
  • Line up the rocks as before.
  • Compare the first two. In a nearby space put the smaller one on the left and the larger one on the right.
  • Compare the next two from the original line. Put them in order nearby.
  • Continue until you have 6 sorted pairs.
  • Now take the leftmost rock from the first pair and compare it to the leftmost rock from the second pair. Put the smaller rock at the left end of a new line. It is the smallest rock of the four.
  • From whichever pair that smaller rock came from, pick up the other rock. Compare it with the one you are still holding.
  • Put the smaller one next to the first, smallest rock.
  • Pick up the fourth rock.
  • Compare it with the one you are still holding. Put these two rocks in order next to the first two. Now you have four sorted rocks.
  • Continue with the next pair of pairs.
  • Continue with the third pair of pairs. Now you have three lines of four sorted rocks.
  • These were merge operations. Perform a merge operation on the first two lines of four. This results in a line of eight sorted rocks, and the other line of four is still there.
  • Merge the line of 8 and the line of 4.
If you were counting comparisons, there were 33. You went through the rocks four times instead of 11. Half the effort, at the cost of a little more complexity of planning. To sort 52 cards this way, you'd wind up making 253 comparisons. That is about 19% of the original effort to sort the deck. The general formula for comparisons is N Log/2(N), so sorting 10,000 items takes about 133,000 comparisons, or 1/376th the effort!

There are clever variations of the shell sort that reduce overhead a little, but that is basically the most efficient sort method. But we were talking about bogosity here. The bubble sort is quite "bogus" compared to the shell sort, which is what my friend meant. Is more bogosity possible?

Certainly. For example, if you play Klondike solitaire with the cards, you are only going to win about one time in four if you don't cheat, so each game you play will have some number of comparisons that is less than 1,326, unless you win, but the comparisons are accompanied by "overhead": stack moves and dealing and so forth that greatly lengthen the time to produce a sorted deck even if you win the first game. A typical "sort" might take four games of average length about 700 or nearly 3,000 total comparisons.

Let's declare that the bubble sort has a bogosity of 1. Then a series of games of Klondike
that leads to a win would have a bogosity of 2. Other solitaire games will also be 2's, though they vary a little in how frequently you win without cheating.

This is not nearly as bogus as it gets: I don't know what number to give it, maybe 100, but there is a "sort" that has maximum bogosity, so far as I know. We can call it the Bogo Sort:
  • Throw the cards across the room.
  • Pick them up and flip the face-down ones face up.
  • Look through the deck to see if they are in order.
  • If not, repeat. (Even if only one card is out of place! No cheating by moving it)
The average number of times you'd have to repeat this operation to achieve a sorted deck is a number with 68 digits (roughly 4E+67). If you could "throw-pick up-check" once per minute, it would take about 1E+62 years (That is a hundred trillion trillion trillion trillion trillion or so). Now that is real bogosity.

Thursday, October 07, 2010

Data modeling is not for humans

kw: computer science, data modeling

I'm taking an afternoon break from the onerous task of writing specifications for a new database and its workflow application. I haven't done data modeling for a few years, and the rust is evident! This kind of work makes more demands on the memory and imagination than anything else I can think of.

I hate to get halfway through a spec document and run into this kind of conundrum: We have two similar items, sets of vocabulary terms. One set is preferred, the other deprecated. One set must be unique, but the other, which embodies synonyms and abbreviations to the first list, need not be. In other words, particularly for the abbreviations, a deprecated term can have more than one interpretation from the preferred list. Wasn't that as clear as mud?!

OK, our database programmer prefers to put both lists into a master vocabulary table with a flag field that indicates which type each term is. The conceptual database looked fine, and even the logical data model seems workable, but going to the physical model, things get creaky. I am not sure a uniqueness constraint can be enforced for one type of term but not the other. I think I'll have to go back and talk everyone into having different data tables for these, which means rewriting about a third of what I've already done. Oh, joy!

Not the first time I've had to do a massive rewrite. Just a pain. Back to the grindstone…

Tuesday, April 21, 2009

It is all just numbers

kw: book reviews, nonfiction, mathematics, computer science, social trends

Do you sit in front of a computer at work? At least part of the time? It is possible, though still a bit clumsy, for your company to track your work habits. If, as I do, you are in a "managed desktop" environment, there could be an app running in the background that logs your activity. While the times of "no activity" may reflect several kinds of "break"—think time, a pit stop, getting a snack or drink, a phone call or colleague's visit—the ratio of active/inactive is just the first ratio a logging program will produce. Perhaps the program also notes time spent in MS Word, Excel, your e-Mailer, surfing with Firefox (and the search terms you enter). Today's PCs have sufficient power to obsessively gather every click and keystroke and when they occurred, while running your applications, seemingly without noticeable impact.

Now consider: if you wear an RFID-tagged badge all day, as I do, it is possible to track your movements about the facility. The PC you are currently not typing at may not "know" why, but a building system could determine whether you really are in the restroom, break room, hallway, or in front of the screen (reading this), and prehaps if the phone is in use. The tag readers that I use to enter my building need the tag to be waved within six inches (15 cm). Who knows whether more capable readers are located throughout the building? I don't. One thing is sure: someday they will be.

Life gets easier once you get home. Off goes the badge. For most of us, the TV doesn't watch us back. But there are test being conducted, particularly for the sake of the elderly, of weight-measuring floor tiles, motion sensors with full-house coverage, perhaps widgets attached to refrigerator, stove and microwave, recording personal habits. When those habits change, or there is a sudden change in weight or in level of activity, someone can be alerted to see how the "subject" is doing. For many frail folk, this could be a godsend. At what point in your life would it go from being "they know too much" to "I need it"?

These are a few instances in which numerical modeling of data being continually gathered is gaining impact in our lives. How soon will it be that the grocery cart in your favorite market will recognize you and show you a route on a GPS-enabled screen, optimizing the route you need to take to fulfill your shopping list—and it already knows what your usual shopping list contains. You could tell it three extra items, and a new route will reflect your new list. It may also lead you past a display or two, calculated to entice you to make an impulse purchase that accords with your (known) tastes. How many people would run screaming from that store, looking for a Mom-N-Pop grocery with "dumb" carts? How many of their children would accept it without even a shrug? Give it time.

In his new book The Numerati, Stephen Baker makes these and other points, though a little less starkly, among his interviews with many of the leaders of the many efforts to turn everything that happens into data, and crunch that data on stacks of supercomputers. When will it be that there is a stack of numbers somewhere, that comprises a mathematical model of Polymath07, knowing my tastes and proclivities, my favorite color (and favorite vice), my favorite everything, my political tastes (or lack thereof), the kinds of people I prefer to be with, my favorite kinds of reading material (a summary of this blog will comprise that part), and whether I prefer having a cat or a dog (I'll never tell; let them find out), or both? What do I collect? Do I have house plants? What kinds? Do I mow my own lawn or hire someone? (Ha! You found out I have a lawn!!).

There was a time that "efficiency experts" such as Frank Gilbreth (of Cheaper by the Dozen) walked around the factory floor with clipboards, optimizing each worker's routines. We will soon be outnumbered by our virtual efficiency experts. Is that a good thing? After chapters showing what the Numerati are working on, in the realms of the workplace, the home, the hospital, the voting booth, national security, even dating services, Baker begins to tell us, Learn to use the numbers for yourself.

We won't all be mathematicians. But we do have some control. Don't like your bank selling your data? How about selling it yourself? From page 205:
One nonprofit organization founded in 2005, AttentionTrust, … provides Web surfers with the tools to amass their own data and sell if, if they choose, to advertisers.
Does that sound like a good idea? It may be. I looked for the organization, but the links are all dead. They may be defunct already. If they are, someone else will probably try soon. What else could one do?

It is hard to say. To date, most people have relied on "security by obscurity." Be ordinary, do ordinary stuff, don't get noticed, and you'll be ignored. Now that the computer power exists to crunch through all the data we all generate, all the time, that is less possible. An idea I like? Be just a little bit subversive. Do you have a "loyalty card" at the supermarket (or several – I carry five)? Get a second one at the same market. Use a variation of your name. Shop differently when you're using the alternate card. Use a different credit card, or cash. When the smart cart arrives, go ahead and follow its path. Then take some of the paths not taken; see what specials are being offered in the other aisles. Think: what does the grocer want from me? Shift your buying pattern to inject confusion into the process. You can extend such ideas into many areas of life.

You may find that adding some randomness and serendipity into your routines will make life more fun! Data-analysis routines are designed to figure out your predictable behaviors. Be less predictable. A side note: In my lifetime, I have taken training classes in several martial arts. The best advice I got was from Joe Begala, the author of the Army's 1940s "hand-to-hand combat" training material. He taught a "Self Defense" class at my college in the 1960s. He said, "If you have to fight, don't stand in any particular way. I can tell if a guy knows Karate, or Aikido, or Boxing, by the way he stands. Then I know something about him. I don't want him to know that about me." I can say this, his class should have been titled "Applied Street Fighting".

Will our grandchildren live in a world not just measured, but controlled, by the Numerati? The Numerati will try. It is up to all of us to keep them off balance.