Tuesday, January 27, 2015

Thoughts from one American mind upon closing "Closing of the American Mind"

I have just finished reading a very strange book, and I want to try to make some personal sense out of it because it was a little unsettling. I'm really not sure what I just read, and I don't think I agree with much of it. But I know that there were some extremely provocative ideas presented here and intellectual honesty compels me to confront them, hence the present post.

The book is titled "The Closing of the American Mind." It was written by Allan Bloom, a UChicago philosopher and classicist who died a few years after its publication in 1987. If you have read Plato's Republic, chances are good you read his translation of it. Anyway, I first noticed TCOTAM this on a shelf in the office of my graduate advisor and was piqued (in both senses) by the title; I saw it again in a box labelled "free books" out in the hall after she had retired, so I took it.

Let me start by saying that this was NOT at all the book I was expecting to read. I was expecting criticisms of higher education, not realizing that this was going to entail questioning the privileged place of freedom, equality, and science provided for by democracy and social contract theory. Thus, by its content alone the book is sure to offend almost everyone. Only about a tenth of it deals with education directly; all the rest is a philosophical critique of modern society. Worse, the author's tone is, well...  imagine your curmudgeonly old neighbor saying "you kids have it too easy these days, with your easy sex and your rock music" and at the same time John Keating, the Robin Williams character in Dead Poets Society, saying,
"Boys, you must strive to find your own voice. Because the longer you wait to begin, the less likely you are to find it at all. ...We don't read and write poetry because it's cute. We read and write poetry because we are members of the human race. And the human race is filled with passion. And medicine, law, business, engineering, these are noble pursuits and necessary to sustain life. But poetry, beauty, romance, love, these are what we stay alive for. To quote from Whitman, "O me! O life!... of the questions of these recurring... carpe diem, seize the day boys, make your lives extraordinary"
I am emphatically not saying that I am persuaded by his arguments (which are often hard to follow and scattered over many pages) or that I agree with him on any of the issues I'm about to raise. He is an idealist and perhaps even a bona fide dualist, but he is nothing if not open-minded; I have tried to approach the book in a similar spirit. I know I have been intrigued and unsettled by his book, largely because it forced me to consider may of the "good" things about our modern world that I had accepted uncritically and now take for granted. This is not to say that he is right about any of this; in fact, I think he is quite wrong about a lot of it. I want to summarize and perhaps explore some of these seemingly heretical challenges, especially the idea that science and democracy, for all they good they have done us, have come to have some serious downsides that require our consideration, not just our blind and passive acceptance.

The Thesis? The Main Idea? Oh, Boy...

The thesis of the Closing, as I have read it, is so knotty that it will take at least a few paragraphs to cobble together some approximation of the thing, and the rest of this blog post to explain what it all means. There are many links in the chain, and while many of them are tenuous, the chain itself is a startling thing that compels inspection. So here goes:

The social contract, which is the foundation of modern democracy, is based on the worst parts of humanity (selfishness, greed, bestial pleasures...). Acceptance of these things as fundamental aspects of the human condition denies us the ability to exercise individual virtue and to experience deep connections to ideas eternity or even to past nobility. It provides for freedom and equality, but these ideals have silently destroyed the possibility of "true religion" and "culture" by putting them all on even footing. Freedom and equality cannot preserve cultures and religions because these latter are based on real differences in fundamental beliefs about good and evil, mutually exclusive claims about what is highest. Thus, there can be no religion; but inasmuch as man needs culture, the deeper spiritual impulse remains. Are we, in fact, happier now? Is this the best of all possible worlds?

Wholesale equality is a happy fiction that kills claims of true goodness, of betterness, in short of any possibility of rank-ordering things. A fundamentally relativistic society has emerged where to discriminate between good and evil is wrong, and therefore everyone is to be indiscriminate. In modernity, we again find ourselves in the state of nature; self-preservation is codified in our laws, but we are still carrying out the old war of all-against-all through a society based on commercial acquisitiveness; further, our increasing openness to everything has led to a hedonism which prematurely and incompletely satisfies our most important human passions and desires, removing their real significance and leaving young people feckless and jaded. The passions should be bound up in the glorious project of fulfilling human nature and searching for final truth, and this binding should be the goal of education.

In a society of rampant freedom and equality whose relativism denies the special claim of reason (for to give priority is anti-equality!), the University provides a necessary counterpoint where unequal things can be pursued by unequal talents; where reason can reign, and where the human spirit can reach great heights in the search for truth and return with the spoils of its labors to benefit humanity at large. Bloom worries that this last bastion of creativity and human exaltation is being threatened.

Bloom asks, "What is man?" Are we, finally, truth-seeking creatures who live for the purpose of discovering the good, or are we comfort-seeking creatures who live only to satisfy animal urges and avoid what is bad? Life presents a chaos of opposites: reason-revelation, freedom-necessity, democracy-aristocracy, good-evil, body-soul, self-other, city-man, eternity-time, being-nothing. He maintains that "a serious life means being fully aware of the alternatives, thinking about them with all the intensity one brings to bear on life-and-death questions, in full recognition that every choice is a great risk with necessary consequences that are hard to bear."

 Vulnerable Man Seeking Means to His Preservation

Bloom begins by noting that the US has one of the longest uninterrupted political traditions of any nation in the world; that is, we have hit upon a good scheme for self-preservation. With the success of political regimes founded on freedom and equality, questions of political principle and of right seem to have been solved once and for all.

Following the teachings of Hobbes, Locke, and Rousseau, the Founders saw that men in the “state of nature” are solitary, selfish savages united only by an animal instinct for self-preservation (i.e., a healthy fear of death). The fear of death is stronger than the desire to dominate others, and so people agree to respect others' lives and property, so long as those others can be forced to reciprocate. "Prudence points towards a good police force to protect men from one another", not toward regimes dedicated to the cultivation of rare and difficult human virtues.

This was actually a pretty radical take on the old political problem; before the social contract, man was thought of as a dual being: one part concerned with the common good and the other part concerned with private interest. For the sake of society, man has to overcome the selfish part of himself in order to be virtuous. A virtue governs a passion, as moderation governs lust or courage governs fear. A man of great virtue was one who was able to conduct himself accordingly. Now, man no longer has to regulate his self-interest; the social contract removes individual virtue by codifying it into law. There's a world of difference between doing something because you personally believe it is right and doing something because external factors force you to do it.

Worse still, self-interest and domination of others still carries the day, only instead of competing for survival and reproduction, violence is now played out in a token economy for status and material wealth. The primacy of self-concern is no longer questioned, much less denied; it is evident in our fundamental institutions, where the original selfishness of the state of nature remains and where concern for the common good is hypocritical. According to Bloom, “Locke had illegitimately selected those parts of man he needed for his social contract and suppressed all the rest, a theoretically unsatisfactory procedure and a practically costly one” (176). For all of its many success, the social contract seems to be based on the worst part of human nature and to eliminate the exercise of individual virtue.


Freedom! Equality! Freedom! Equality!

In America, consent of the governed is based on freedom and equality. But does "equality" mean only the equal opportunity for unequal talents to acquire property? Should these mercenary skills be better rewarded than moral goodness or other high human traits? Does "freedom" simply mean the ability to exercise our baser impulses at leisure, provided they don't harm others? These ideals have assumed a sacred status in our society; they are, like religions of old, beyond the reach of criticism and not up for discussion. I certainly had never thought to cast a critical eye on them before; they seem to be "good", certainly better than most everything that came before them.

If freedom is not a natural right, it is at least a very insidious meme (or self-perpetuating cultural phenomenon). The American Founders wanted to mitigate extreme beliefs, particularly religious beliefs, because these inevitably lead to civil strife. So all religious sects have to obey a common law and be loyal to the Constitution, and if they do so, everyone else had to leave them alone, regardless of their beliefs. “Freedom of religion,” however, is based on the sneaky premise that there is no single true religion and everyone is entitled to practice their own. Thus, Religion belongs to the realm of opinion and choice, not the realm of knowledge and truth. Freedom itself, however, is no matter of opinion; it is a fundamental right and an ultimate Truth. You are free do anything except challenge freedom; all ideas have equal status, except equality. Bloom notes, “The most successful tyranny is not the one that uses force to assure uniformity, but the one that removes the awareness of other possibilities, that makes it seem inconceivable that other ways are viable, that removes the sense that there is an outside” (249).

Here again we live with two contradictory understandings of what counts for man. One tells us that what is important is what all men have in common; the other that what men have in common is low, while what they have from separate cultures gives them their depth and their interest...Human rights are connected with one school, respect for cultures with the other. Sometimes the United States is attacked for failing to promote human rights; sometimes for wanting to impose "the American way of life" on all people without respect for their cultures. To the extent that it does the latter, the US does so in the name of self-evident truths that apply to the good of all men. But its critics argue that there are no such truths, that they are prejudices of American culture.  (191)
Privileging freedom and equality above all else means that we are not permitted to seek natural human greatness and admire it when it is found, nor permitted to recognize badness and have real contempt for it. Everyone is entitled to their own, and no one's is any better than another's.

The Founders constructed this elaborate political machinery in order to contain minority beliefs and factions in such a way that they would cancel each other out and allow for the pursuit of the common good. The goal was to achieve a national majority concerning the fundamental rights, and once achieved, to prevent that majority from using its power to overturn those rights. That dominant majority gave the country a dominant culture with its own traditions and special claims to knowledge and taste.

The Constitution guarantees protection for the rights of individual human beings, irrespective of race or creed. It does not promise respect for minority belief, but merely toleration. But from the very beginning, there have been individual factions demanding unique respect for their own unique beliefs and ways of life. To these groups, the founding principles of freedom and equality were impediments, and they tried to overcome the principle of majoritarianism in favor of a nation of minority groups each pursuing their own interests.

An unexpected early example of the challenge of majoritarianism was the South and the question of slavery. The Constitution was clearly a document with a moral commitment to equality and hence condemned slavery. Yet southerners were quite successful at characterizing their "peculiar institution" as part of a charming diversity and individuality of culture which was being threatened. Openness and relativism was just what was needed to defend their way of life against the intrusions of others who called for equal rights. 

Bloom's final argument against openness and relativism is this:

Herodotus was aware of the rich diversity of cultures... but he took that observation to be an invitation to investigate all of them to see what was good and bad from them. Modern relativists take that same observation as proof that all such investigation is impossible and that we must be respectful to them all. ...I know that men are likely to bring what are only their prejudices to the judgment of alien peoples. Avoiding that is one of the main purposes of education. But trying to prevent it by removing the authority of men's reason is to render ineffective the instrument that can correct their prejudices. True openness is the accompaniment of the desire to know, hence of the awareness of ignorance. To deny the possibility of knowing good and bad is to suppress true openness.


Nihilism, Science, and Human Dignity.

I am, for the record, a flaming reductionist and a life-long atheist; I believe that through natural science, everything is explicable in terms that we can finally understand and that the realities of life and nature are accessible to us if we can discover them. Things like "dualism" and "soul" are anathema to my worldview. What this book has impressed upon me is that that, while I still hold fast to this belief and have amassed much convincing evidence that I am correct, it is still my particular belief -- my vision of the way things really are -- and it may, in the last analysis, be a bad choice of polestar by which to guide a human life to say nothing of its truth-value. After all, my beliefs, values, and allegiances largely reflect those of middle-class American society circa 1990 onward; in the whole belief/value possibility space, ours is almost certainly NOT a global optimum. (Do we even know what to maximize? Happiness?)

Science is fundamentally reductionistic and deterministic; it denies any special place for human dignity. Humans appear to have some kind of special status, at least in their capacity to reason, but their capacity to reason has left them out in the cold. Bloom poignantly writes,
"God is dead, Nietzsche proclaimed. But he did not say this on a note of triumph, in the style of earlier atheism---the tyrant has been overthrown and man is now free. Rather he said it in the anguished tones of the most powerful and delicate piety deprived of its proper object. Man, who loved and needed God, has lost his Father and Savior without possibility of resurrection." (195)

Reason has discerned that all previous cultures were founded by gods or belief in gods. Only if our new Enlightened regimes are extremely successful, able to rival the creative genius and splendor of these cultures, could reason's rational foundations be equal or superior to the kinds of foundings that reason knows were made elsewhere. Since this superiority is questionable, reason recognizes its own inadequacy. Nihilism seems to be the call to abandon reason on rational grounds; the old worlds of meaning have crumbled in the light of truth. Man as creator can overcome this void, but only by accepting checks-and-balances on freedom and equality:
Nietzsche... held that inequality among men is proved by the fact that there is no common experience accessible in principle to all... The individual value of one man becomes the polestar for many others whose own experience provides them with no guidance. The rarest of men is a creator, and all other men need and follow him...
Authentic values are those by which a life can be lived, which can form a people that produces great deeds and thoughts. Moses, Jesus, Homer, Buddha: these are the creators, the men who formed horizons, the founders of Jewish, Christian, Greek, and Indian culture. It is not the truth of their thought that distinguished them, but its capacity to generate culture. A value is only a value if it is life-preserving and life-enhancing. The quasi-totality of men's values consists of more or less pale carbon copies of the originator's values. Egalitarianism means conformism, because it gives power to the sterile who can only make use of old values, other men's ready-made values, which are not alive and to which their promoters are not committed. Egalitarianism is founded on reason, which denies creativity.
He goes on to argue (I think) that, if all of this is taken as true, then our regime --founded on reason and dedicated to the protection of freedom and equality above all-- is doomed. I am a little out of my league talking about this stuff, and I am not particularly sympathetic to it, so I'll leave it at that. It is not particularly convincing to me, and it's probably not convincing to you either. However, I will close this section with a single passage that I think begins to illustrate the point he wants to make, and in which I see a glimmer of truth:

My grandparents were ignorant people by our standards, and my grandfather held only lowly jobs. But their home was spiritually rich because all the things done in it, not only what was specifically ritual, found their origin in the Bible's commandments, and their explanation in the Bible's stories and the commentaries on them, and had their imaginative counterparts in the deeds of myriad exemplary heroes. My grandparents found reasons for the existence of their family and the fulfilment of their duties in serious writings, and they interpreted their special sufferings with respect to a great and ennobling past. Their simple faith and practices linked them to great scholars and thinkers who dealt with the same material, not from outside or from an alien perspective, but believing as they did, while simply going deeper and providing guidance. There was a respect for real learning, because it had a felt connection with their lives. This is what a community and a history mean, a common experience inviting high and low into a single body of belief.

I do not believe that my generation, my cousins who have been educated in the American way, all of whom are MDs or PhDs, have any comparable learning...I am not saying anything so trite as that life is fuller when people have myths to live by. I mean rather that a life based on the Book is closer to the truth, that it provides the material for deeper research in and access to the real nature of things. Without the great revelations, epics, and philosophies as part of our natural vision, there is nothing to see out there, and eventually little left inside. The Bible is not the only means to furnish a mind, but without a book of similar gravity, read with the gravity of the potential believer, it will remain unfurnished. (60)

 Education? Isn't This a Book About Education?

If I agree with nothing else Bloom says, I can at least give my unqualified support to his central message about education: students learn best when they are passionate about their learning. Here's something I can really get behind!

He finds, however, that the relativism and openness of society has sapped students' curiosity about final things. Everything's settled, isn't it? When your truth is as good as mine, when your good is as true as mine, what's left to think about? And since our social status quo gives selfishness free reign (e.g., capitalism), we have thereby accepted its primacy in human nature. Boom, done. People are selfish and there is no truth, case closed. All that's left is to wage a status war with our fellow man, the spoils of which allow us to find ever more exotic ways of satisfying our lowest instincts (hunger, sex, comfort). This sort of jaded thinking, Bloom argues, is relatively new and has some upsetting consequences for human nature. Taking the pulse of the student body, Bloom would ask his class questions like “What books really count for you? Who are your heroes? Who do you think is evil?” Most have no answer.

Bloom thinks Idealism should have primacy in education, for we must take our orientation by our possible perfection. "Utopianism is, as Plato taught us at the outset, the fire with which we must play because it is the only way we can find out what we are." Deprived of literary guidance and its examples of high human types, students no longer have any image of a perfect soul, and hence do not long to have one. They do not even imagine that there is such a thing. And neither do they have any idea of evil; they doubt its existence. Hitler is just another abstraction, an item to fill up an empty category. "Bloom argues that "the most common student view lacks an awareness of the depths as well as the heights, and hence lacks gravity" (67).

But why? Easy satisfaction of basic pleasures make material desire the proper goal of a life. Bloom maintains that the premature cheapening of human passions "ruins the imagination of young people and makes it very difficult for them to have a passionate relationship to the art and thought that are the substance of liberal education." Today's youth experience premature ecstasy through casual sex and drug use. These things, he argues, "artificially induce the exaltation naturally attached to the completion of the greatest endeavors--victory in a just war, consummated love, artistic creation, religious devotion, and discovery of truth. Without effort, without talent, without virtue, without exercise of the faculties, anyone and everyone is according the equal right to the enjoyment of their fruits." (80)

Education is not sermonizing to children against their instincts and pleasures, but providing a natural continuity between what they feel and what they can and should be. By giving kids nothing to believe in, we give them no basis for discontent with their understanding of the world and no awareness of alternatives. They are made complacent by an acceptance of self-interest as the defining feature of humanity (perhaps it is?), leading to a culture of all-relativism and instant gratification. Who needs any more meaning than this?


Friday, December 19, 2014

"The Greatest Equation of All Time"

Sorry to have been so long away from the blog! The purpose of this post is to give some intuition for Euler's identity, "the most beautiful theorem in mathematics", to those who haven't seen it before or to those for whom it is meaningless because they lack the mathematical background.

It was never introduced to me in any class, and unless you majored in math or physics the same is probably true for you. I wanted not only to call your attention to this equation, but also to derive it, the derivation of which is equally marvellous. Walking through the derivation with a pencil was perhaps my first real taste of what I imagine mathematicians live and breathe for; my present aim is to give you a taste!

Here it is:

It's pretty astounding when you look at it for the first time... it seems a ridiculous assertion! (However, Gauss is supposed to have said that if this was not immediately apparent to you on being told it, you would never be a first-class mathematician.) He also had a really cool signature, so he is probably right.

Really though? Are we talking about e (2.71828...), like from continuously compounded interest? You mean π (3.14159...), like a circle's circumference-to-radius ratio? And the imaginary number i, the square root of -1, that we talked about for like a week in high school and then never actually used for anything? Yep, those are the ones, e, π, and i. Dust off your old TI-8x, punch this badboy in, and see what output you get. The batteries are probably dead so I'll spare you the trouble:


Though I am mainly interested in showing you a way to get to this shocking identity, I will need thereby to talk a bit about Taylor series approximations. If you don't know much about this technique, think of using a sum of polynomials to approximate the shape of function at a given point. If you don't know much about polynomials and functions, think of adding together easy-to-compute things (x, x2, x3, ...) to approximate hard-to-compute things (ex, sin(x), ln(x), ...). This is actually how your calculator computes with things like e, π, et al. 

But this isn't a post about Taylor series approximations, so I'll cut down on the details. If you already understand how to use Taylor series approximations, you could skip to the derivation. If you aren't, it is super useful and cool, so check it out! First things first: if you want to approximate the shape of a function f(x) using an easier-to-work-with polynomial like a power series (i.e., a + bx + cx2 + dx3 ... where a, b, c, d,... are constants), you can do it like this:



where f(n)(0) represents the nth derivative of function f(x), evaluated at x=0 (this gives us the shape around x=0) Technically, this is a Maclaurin series. Really cool! Because if we can keep taking derivatives, we can get an almost perfect approximation of f(x) using polynomials instead. Don't worry, I'm not going to ask you to do any calculus (though you might want a helpful refresher on derivatives*). I'll just show you the results.

To follow the derivation, you just need the following 3 results: the derivative of sin(x) is cos(x), the derivative of cos(x) is -sin(x), and the derivative of ex is itself, just ex.


Let's do sin(x) first, because I've got some pretty pictures of it!



We want to approximate this function using a Taylor series polynomial. Let us do so by taking the first term of that big equation above:



So using only 1 term, our approximation is the line y=x


Not a very satisfying approximation, is it? What about 2 terms?




If we plot x - x3/3!, we get


Hey wow, that's a lot better! The polynomial x - x3/3! seems to pretty exactly mimic sin(x), at least around x=0. Let's add another term.


Now our approximation is y = x - x3/3! + x5/5! which looks like this:



Even better! You can see that by adding more terms of the Taylor polynomial, we get a closer approximation of the original function. Here's a couple more terms just to illustrate:




Now the polynomial approximation is virtually indistinguishable from sin(x), at least at from x=-pi to x=pi. Here's a more "moving" display of this (credit). Notice how adding more and more terms gets us an ever-better approximation of sin(x).



Here's an even easier example! (We are going to need this result too.)




So we have Taylor series approximations for sin(x) and for ex. Now all we need is one for cos(x). It is just like sin(x) except all of the opposite terms cancel out! I'll leave this as an exercise to the reader (always wanted to say that!) and just give you the result:


Take a minute to look at these. Strange, isn't it? If it weren't for those negative signs, it looks like simply adding sin(x) and cos(x) would give us ex... Hmmm.


We are now in a position to make some magic happen, but one more thing remains to be done. Now's the time to recall your imaginary number i, which is equal to −1. Since i = −1, we know i2 = -1. For higher powers of i, we have the following pattern:


The pattern i, -1, -i, 1, repeats as you take ever higher powers of i. Notice that the signs are switching too! You should be excited about this. Now we're ready for the main event!



YEAH! WOOHOO! At least, that's how I felt when I first saw it.

And it's not just superficially beautiful either. Before we plugged in pi, we saw that eix=cos(x) + i*sin(x). This formula has many important applications, including being crucial to Fourier analysis. For an excellent introduction, check this out.

Well, I hope that if the derivation didn't leave you reeling in perfect wonderment, you were at least given something to think about! Thanks for reading!
________________________________________________________________
*Quick derivatives refresher
It may be helpful to think of the derivative of a function f(x)---symbolized as d/dx f(x) or f'(x)---as a machine that gives of the slopes of tangent lines anywhere along the original function: if you plug x=3 into f'(x), what you get out is the slope of a line tangent to the original function f(x) at x=3. Since slope is just rise-over-run, the rate at which a function is increasing or decreasing, the derivative gives as the rate at which the original function is increasing or decreasing at a single point. If f(x) is the parabola x2, then its derivative f'(x) is 2x. At x=0, the very bottom of the parabola, we get f'(0)=2(0)=0, which tells us the line tangent to x2 at the point x=0 has zero slope (it's just a horizontal line). At x=1, the parabola has started to increase; the rate it is increasing at that point (the slope of the line tangent to that point) is f'(1)=2(1)=2; so now we have a line that goes up 2 for every 1 it goes over. At f'(2)=2(2)=4, a line that goes up 4 for every 1 it goes over. This agrees with our intuition when we look at a parabola; it is accelerating upward at an ever increasing rate.

Thursday, December 18, 2014

"Important Peculiarities" of Memory

In my high school psychology class I was told that human memory capacity was unlimited ...and it has bothered me ever since. How, I mean? Aside from the physical limitations on information storage, how could a system that remembers everything forever be evolutionarily advantageous?

This is a question I hope to explore in a deeper way sometime soon; for now, I want to talk to you about a few "peculiarities of human memory" that begin to shed some light on the situation (Bjork & Bjork, 1992). Know that I am drawing heavily from this source and their Theory of Disuse for the present discussion. This is really the coolest part, but I've left it until the end. First, let's talk about three "peculiarities"...

1. STORAGE AND RETRIEVAL ARE TWO DIFFERENT THINGS:
Analogies of human memory -- to a bucket being filled, to computer memory, to magnetic tape -- are often grossly misleading. No literal copy is recorded when you store a piece of information in memory. Learning isn't opening a drawer and putting something in; remembering isn't opening a drawer and taking something out. Indeed, your brain is not a drawer.

New things are placed into memory via their semantic connections to things already in long-term memory. The more knowledge you have of a given area, the more ways you have to store additional information about it. This is a strange biological instantiation of the Matthew effect, where "the rich get richer". To me, one of the most incredible things about being a living, thinking human is this virtually unlimited capacity for storing new information. And when I say "incredible," I mean it both in the sense of "wow, golly!" as well as in the more literal sense of "not credible."

"But wait," I hear you ask, "if my memory is so bally infinite, why can't I remember my passwords half the time? And I'm always forgetting peoples' names, and I can't remember a single word of that book I read last week, and..." As it turns out, getting information into memory is easy, but getting it out is quite another matter.

Quick, what was your childhood address? Your first cellphone number? Your seventh grade math teacher's name? Your high school ID card number? How about your old AIM password? Even the most repetitively drilled, frequently accessed pieces of information eventually become inaccessible through years of disuse. Weirdly though, this information is still stored in memory: you could probably correctly identify each of the above from a list of distracters, for example, and you probably wouldn't have any trouble remembering if you were back in the context of your home town. Perhaps if I had asked you on a different day, when you were in a different mood or frame of mind, you would've been able to retrieve the information. Often information that is effortlessly recallable on one occasion can be impossible to recall on another. Maybe you weren't able to muster the answers at first, but now after expending a bit of time and effort you have remembered. Should the old information become pertinent again, it will certainly be relearnable at an accelerated rate.

What we can and cannot retrieve from memory at any given time appears to be a function of the cues that are available to us at that time. These "cues" may be general situational factors (environmental, social, emotional, physical) as well as those having a direct relationship to the to-be-retrieved item. Cues that were originally associated in storage with the target item need to be reinstated (physically, mentally, or both) at the time of retrieval.

The main takeaway here is that our capacity for storage is far greater than our capacity for retrieval, and it appears that fundamentally different processes are responsible for each. Storage strength represents how well an item is learned, whereas retrieval strength indexes the current ease of access to the item in memory. These two strengths are independent: Items with high retrieval strength can have low storage strength (e.g., your room number during a 5-day stay at a hotel).

2. RETRIEVAL MODIFIES THE MEMORY!
The mechanical analogies have other flaws; reading from computer memory does not alter the contents, whereas the act of retrieving information in human memory modifies the system. When you remember something, that piece of information becomes easier to remember in the future (and other information becomes less retrievable). This is why taking tests is better for long-term retention than studying is. More odd is the idea that recalling Thing A can make it more difficult to recall Thing B in the future, an effect sometimes known as "retrieval competition." This topic is very, very interesting but that's all I'm going to say about it here. For more on the testing effect, check out this paper.


3. LONG-HELD MEMORIES ARE HARD TO REPLACE
Through disuse, then, things become hard to retrieve. But oddly, the earlier the memory was constructed (i.e., the older it is), the more easy it is to access relative to related memories constructed later. Say you decide to change your email password; after doing so, the new password will be the most readily accessible of the two. If you use it to log in tomorrow, you will have little trouble recalling it. However, if you do not have occasion to use either password for a while (new or old), the old password becomes far easier to remember relative to the new password.

Consider athletes: a long layoff often leads to the recovery of old habits. This can help an athlete recover from a recent funk, or it can be a major setback for a rookie who has been rapidly improving. In occupational settings too, and even the armed services, people can appear to be well-trained but then turn around and take inappropriate actions at a later time (i.e., fall back on old habits), particularly in stressful situations. It may even result in the unreasonable surprise we often feel when we see that a child has grown, or a friend has aged, or a town has changed; perhaps we are overestimating these changes because our memory of the child, friend, or town is biased toward a past version of them stored more securely in memory.

This stuff has firm support from laboratory studies too, but I don't want to bore with great detail. Suffice it to say that, if experimental subjects are given a long list of items to memorize and then afterwards asked to recall all of the times that they can remember, there will be a strong recency effect: the items later in the list will be more easily recalled. If, later on (say a day or a week later), subjects are asked to recall all of the items they can remember, there will be a strong primacy effect: the items appearing first will be better recalled than the items appearing later. That this, with the passage of time there is a change from recency to primacy. This finding holds across different delays, tasks, materials, and even species (Wright 1989)!


The Theory of Disuse:
In brief, Bjork and Bjork's (1992) theory states that items of information, no matter how retrievable or well-stored they are in memory, will eventually become nonrecallable if they are not used enough. This is not to say that the memory has decayed or been deleted... it is just inaccessible. Storage and retrieval are two very different things; storage strength reflects the how well-learned the item is, while retrieval strength represents how easy it is to access the item. Unlike storage capacity, retrieval capacity is limited; that is, there are only so many items that are retrievable at any given time in response to a cue or set of cues. As new items are learned, or as the retrieval strength of certain items in memory are increased, other items become less recallable. These competitive effects are generally determined by category relationships defined semantically or episodically; that is, a given retrieval cue (or cues) will define a set of associated items in memory, and the dynamics of competition for retrieval capacity take place across the set.

The theory makes many predictions that account for those peculiarities stated above. Retrieval capacity is limited, not storage capacity, and the loss of retrieval access is not a consequence of the passage of time per se, but of the learning and practice of other items. Retrieving an item from memory makes it easier to retrieve that item in the future but makes it more difficult to retrieve other associated items. The theory also explains why overlearning (additional learning practice after perfect performance is achieved) slows the rate of subsequent forgetting: perfect performance is a function of retrieval strength (which cannot go beyond 100%), whereas additional learning practice continues to increase storage strength. Finally, the spacing effect--the fact that spreading out your study sessions is far more effective for long-term retention than is cramming--can be accounted for by the theory as well. Spacing out repetitions increases storage strength to a greater extent than does cramming, which in turn slows the rate of loss of retrieval strength, thereby enhancing long-term performance. Importantly, cramming can still produce a higher level of initial recall than that produced by spacing, but like the switch from recency to primacy, the switch happens rather quickly.

Again, to spin an evolutionary just-so story, all of this seems pretty adaptive. It is sensible that the items in memory we have been using lately are the ones that are most readily accessible; the items that have been retrieved in the recent past are those most relevant to our current situation, interests, problems, goals... and in general, those items will be relevant to the near future as well. To keep the system current, it makes good sense that we lose access to information that we have quit using: for example, when telling someone our address, it would not be useful to recall every home address we have had in the past.

I look at all of this and I see a selection process at work. The set of items in memory is like so many species in an ecosystem; introduce a new species, and it will die unless it finds a niche (new information must be learned well enough to make it into long-term memory in the first place). Some species don't have much to do with one another, whereas others are mutually dependent and others in direct competition (increasing the retrieval strength of one item reduces the retrieval strength of other, related items). Species with low fitness diminish relative to those with high fitness because they cannot stay  competitive (the items that are used the most proliferate at the expense of items that don't, the item's fitness being determined by the history and recency of its use). Longer, more established species are better adapted to their environment and thus tend to outcompete newcomers (older, more well-connected items memory are easier to recall than newly learned items lacking a deep connection to other items in memory). Species die out, but rarely go completely extinct; instead, they can emigrate elsewhere. They are still extant, but no longer part of the active ecosystem. When conditions improve and it becomes adaptive to return to the ecosystem again, the species is easily reinstated. My metaphor falls apart in places, but I find the selection scheme a good jumping of point for most discussions of this nature.


Thursday, October 30, 2014

2 x 2 Statistics


Modern data analysis has gotten very complicated! Let's forget all that for a moment. I wanted to take this opportunity to examine several useful statistical techniques and measures of association that involve nothing more than a 2 x 2 contingency table. A 2 x 2 contingency table (also called a cross tabulation or cross tab) is a simple grid that displays the frequency distribution given two variables--one variable for the columnsand one variable for the rows. Suppose, for instance, that one variable is sex (male or female) and one variable is handedness (right- or left-handed). Lets say we took a random sample of 100 people and determined both their sex and their handedness. We could then take this data and create a contingency table. There are 2 possibilities for sex and 2 possibilities for handedness, resulting in 4 unique combinations: right-handed males, right-handed females, left-handed males, and left-handed females:


 Right-handed   Left handed  total:
Male  43952
Female 44448
total:8713100

Here, if the proportions of individuals in different columns vary significantly between rows (or if the proportions in different rows vary between columns), we say there is a "contingency" between the two variables, and thus they are not independent of each other. This is multivariate statistics in its most basic, two-variable form, but many useful analyses arise from it. Chances are good that you've heard of most of them, and you may have even had occasion to use a few of them before! They are often spotted on the fringes of SPSS output but rarely ever talked about. But many can be quite useful! I'm going to talk about Chi-squared tests, Fisher-exact tests, odds ratios, and several approaches to the humble correlation coefficient (Phi, Rho, biserial, point biserial). Take a look!

Chi-Squared Test





Q1 Right Q1 Wrong total:
Q2 Wrong a b a+b
Q2 Right c d c+d
total: a+c b+d a+b+c+d=N


The above table provides a framework for the examples here to come. Let's say we give a test consisting of just 2 questions (Q1 and Q2) to a total of N students: in the table above,  a is the number who got Q1 right and Q2 wrong, b is the number who got both Q1 and Q2 wrong, c is the number who got both Q1 and Q2 right, and d is the number who got Q1 wrong and Q2 right. Thus, a+b+c+d = N. Also note that a+c is the total number of students who got Q1 right, b+d is the total number who got Q1 wrong, a+b is the total getting Q2 wrong, and c+d is the total getting Q2 right. These numbers are found along the margin of the contingency table.

Now, let's say we were interested to know whether students who got Q1 correct were more likely to get Q2 correct; that is, we want to know if scores on Q1 are correlated with scores on Q2. Let's say we give the test to N=100 students, and here are their scores:




Q1 RightQ1 Wrongtotal:
Q2 Wrong143549
Q2 Right411051
total:5545100

 To assess whether or not there is a relationship between scores on Q1 and Q2, we are essentially asking whether, within each column, the rows differ significantly, or vice versa: within each row, do the columns differ? Well, what would it look like if there were no differences?


Q1 RightQ1 Wrongtotal:
Q2 Wrong?...49
Q2 Right......51
total:5545100


Start at the top rightmost cell (Q2 wrong, Q1 right). Assuming that we know only the totals, we would expect that since 49/100 (49%) of students got Q2 wrong, and since 55/100 (55%) of students got Q1 right, then 49% of 55% of the 100 students got Q2 wrong and Q1 right. Thus .49*.55*100 = 27 students. Using the totals, we can subtract 27 to calculate all of the other frequency values:


Q1 RightQ1 Wrongtotal:
Q2 Wrong2722 49
Q2 Right28 23 51
total:5545100

 These are the frequencies we would expect to get if Q1 and Q2 were independent of each other (that is, uncorrelated). But this is not what we saw!

To see if these expected frequencies are significantly different from the numbers we observed, we compute a chi-squared test statistic like so:

Thus, for each cell, we subtract the expected frequency from the observed frequency, square that quantity, and divide by the expected frequency. Then we add them all up. This gives us a measure of how much our sample deviates from the expectation.


Our test statistic is 27.3; this seems big, but we need to look at a chi-squared distribution with (#rows-1)*(#cols-1) degrees of freedom. Here, that's just 1.
Instead of going to a table of critical values, I'll use R to make a pretty graph:

 > qchisq(.95,1)  
 [1] 3.841459  
 > cord.x<-c(0, seq(0.001,3.84,.01),3.84)  
 > cord.y<-c(0,dchisq(seq(0.001,3.84,.01),1),0)  
 > curve(dchisq(x,1),xlim=c(0,10),main='Chi-Squared, df=1, a=.05')  
 > polygon(cord.x,cord.y,col="skyblue")  


If the data were independent, there is a less than 5% chance of observing a test statistic larger than 3.84; our test statistic was much larger, 27.3! Thus, it is extremely unlikely that we would observe a value that large if our data were uncorrelated. So, we reject that possibility.

Another, perhaps quicker way of computing the chi-squared test statistic is worth mentioning here.

Check out the squared quantity on top: look familiar? You've got a determinant here! If that quantity (ad-bc) were zero, you would know that responses to the two questions were perfectly independent. It would also be the case that the 2x2 table is "rank one", meaning that all rows are scalar multiples of the first nonzero row. To illustrate this briefly,


Q1 RightQ1 Wrongtotal:
Q2 Wrong437
Q2 Right8614
total:12921

Multiply the top row by 2 to get the bottom row, right? Also ac-bd = 24-24=0. So here there is no association between scores on Q1 and on Q2.

But wait, back to our original example.

Measuring the extent of the relationship with correlations:

 Phi coefficient

Once you have your chi-squared test statistic, and you know that there is a relationship between your two categorical variables, you can easily use this number to calculate phi, which results in the same value as a correlation coefficient. If we take our chi-squared test statistic (27.3 above), divide it by the total number of people, and take the square root of that quantity, we get  the phi coefficient:

Doing this here yields sqrt(27.3/100)=0.522, which is a moderate correlation. Note that this is exactly the same value as would Pearson's r. Here's a helpful table to help you pick the measure of association that's right for you!


ContinuousOrdinalCategorical
ContinuousPearson's r--
OrdinalBiserial rSpearman's ρ; polychoric-
CategoricalPoint-biserial r Rank BiserialPhi (φ)/Matthews

Odds Ratios


Q1 RightQ1 Wrongtotal:
Q2 Wrong143549
Q2 Right411051
total:5545100


An odds ratio is another way of quantifying the size of the effect detected by chi-square significance tests of dichotomous variables. Think of the "odds" of an event as the probability of the event happening divided by the probability of it not happening. If you got Q1 right (first column), the odds of a correct answer on Q2 are (41 right)/(14 wrong), or 2.93, meaning that almost 3 times as many people got it right as got it wrong, so getting it correct is 2.93 times as likely (the odds are 2.93 to 1). We can calculate other three 'odds' in the same manner; of the people who got Q1 wrong (second column): 35 got Q2 right and 10 got Q2 wrong, so the odds of a right answer to Q2 given you got Q1 wrong are (10 right)/(35 wrong)=.29; thus, if you get Q1 wrong, you are more than three times as likely to get Q2 wrong than you are to get Q2 right.

Now, the odds ratio is just what it sounds like: the odds of an event happening in one group divided by the odds of it happening in another group. The ratio of two odds tells us how much more likely an event is to happen in one group versus another. Using the calculations above, we can ask, "how much more likely are the students who got Q1 right than those who got Q1 wrong to get Q2 right ?" To do this, divide the first odds of a correct Q2 of the first group (those who got Q1 right) by the odds of a correct Q2 of the second group (those who got Q1 wrong). Doing so yields (2.93/0.29)=10.1, which means that those who got Q1 right are 10 times more likely to get Q2 right than are those who got Q1 wrong.
If an odds ratio is greater than 1, then "A" is associated with "B" in the sense that having "A" raises the odds of having "B" (relative to not having "A"). In the present example, "Q1 Right" raises the odds of "Q2 Right" (relative to "Q1 Wrong").


Fisher's exact test

OK, quick aside. Imagine you have an urn containing 10 balls; 3 of them are black and 7 of them are red. If you choose 5 balls at random, what is the probability that all 5 are red?  If you know a little about combinatorics, you know that there are "10 choose 5", or 252 ways to choose 5 balls from 10:

 

This tells us that there are 252 unique 5-ball combinations; now how many of those have 5 red balls and 0 red balls? Well, how many such combinations are there? How many ways are there to choose 5 red balls from 7 red balls? Say "7 choose 5". Now, how many ways are there to choose 0 black balls from 3 black balls? "3 choose 0." The product of these two combinations (7 C 5) and (3 C 0) give us the number of possible 5-ball combinations in which all 5 are red (and zero are black). 



Now then, these what proportion of the total number of possible 5-ball draws (total=252) will have 5 red and 3 black (total=21)? Well,

 

Where N is the total number of balls in the urn, R is the total number of red balls in the urn, n is the total number we select, N-R is the number of black balls in the urn, r is the number of red balls in our selection, and n-r is the number of black balls in our selection. Experiments of this kind follow a hypergeometric distribution, we can this method to analyze our 2 x 2 contingency table.
Now, instead of balls and urns, we have right and wrong. Instead of asking, "what is the probability of getting 5 red balls and 0 black balls when I draw 5 balls at random from a population that contains 7 red balls and 3 black balls" we ask "what is the probability that 14 people got Q1 right/ Q2 wrong and 35 people got Q1 wrong/ Q2 right if we draw 49 balls at random from a population where 55 got Q1 right and 45 got Q1 wrong." That is, we want the probability of observing the given set of frequencies 14, 35, 41, 10 in a 2x2 contingency table given fixed row and column marginal totals.

Q1 RightQ1 Wrongtotal:
Q2 Wrong143549
Q2 Right411051
total:5545100

 

This probability is vanishingly small, but what we are really interested in is this: given our marginal frequencies, what is the chance of observing a difference at least this extreme? More extreme configurations can be generated by locating the smallest (or largest) frequency in the table, subtracting 1 (or adding 1), and then computing the remaining items given the observed marginal frequencies. When you find the collective probability of observing a combination of frequencies at least as extreme as this one, you get something smaller than 1/10,000,000. Thus, we can conclude that answering question 1 right is associated with answering question 2 right.

Fisher's exact test is interesting because it depends on no parametric assumptions and works for any sample size.


I may add more to this post in the future, but that's enough for now!

Sunday, July 13, 2014

The Math of Secrecy: RSA Cryptography (and Shapes You Can Draw?)

When Gauss was 19, he discovered that of the infinite number of polygons that have a prime number of sides, a mere five of them can be constructed with a ruler and compass (i.e., using only straight lines and circles). These prime-sided polygons can have 3, 5, 17, 257, or 65537 sides, but only these five are possible (probably).


Indeed, the only shapes you can draw with an odd number of equal-length sides are the multiples of these 5 Fermat primes: 3, 5, 15, 17, 51, 85,..., 4294967295
(amazingly, all constructable polygons must have a number of sides that is a multiple of a Fermat prime and a power of 2!). There are good reasons why this is true, but they are confusing and would belabor this post.

Apparently, Gauss was so happy with his finding that he requested a regular heptadecagon on his tombstone; the stonemason declined, stating that it would essentially look like a circle. After watching the requisite process below, one begins to sympathize with the stonemason (diligent wikipedian Aldoaldoz made both of these; in case you want to build your own 17-gon, you can break out the compass and follow along at home). It's hypnotically beautiful, but the .gif alone is 462 frames and takes a full 1:26 to watch...


These strange numbers are known as Fermat primes, and take the form 22n+1. Though there are infinitely many numbers of this form, only the first 5 (above) are prime. As I was reading about these yesterday, I found that they have an important application in the most common form of public-key cryptosystems, whereby messages are encoded so as to conceal their meaning.

This method of encryption is used to secure electronic communication over the internet; even if a third party somehow manages to nab your encrypted message, it will be almost impossible for them to crack it. If you've used SSH, SSL/HTTPS, PGP, or had to verify a digital signature, you've used RSA encryption or one of its descendants.

On personal websites, you may have run into these before:

-----BEGIN PGP PUBLIC KEY BLOCK-----
Version: GnuPG v1.4.11 (GNU/Linux)

mQENBFPCyXEBCACuHid62W3FI3DegXw3G6Xyjdj3SBl3+f/fBNIN4Yrx0auPjuZG
TqtA6opOH7jzAEBdBBysiQ+1frQlfiWlmdzJ/GQR7KGhuZNx33pyCwXV85bcKtno
A4CQK8r2sfrRF796voNWxW/MaStT7IWQfHrMYsgcl+7cZogBu/nl3nnHuZz+oMMG
ZZl+uziKF1+M4naOr6gH3UMTECk2Xib2lk58RFN4pmqPzbWG5gUU5ugN13c6hO7S
eKN/cbGSHRHPQci0aZo743rIoWgQZ+S88j3BweGFbD78tw5UYJUW+rnyYISzDbVi
R+i8luzVtVhkHZnetcQoz6IBsDyfnK0dKMLhABEBAAG0K05hdGhhbmllbCBSYWxl
eSA8bmF0aGFuaWVsLnJhbGV5QGdtYWlsLmNvbT6JATgEEwECACIFAlPCyXECGwMG
CwkIBwMCBhUIAgkKCwQWAgMBAh4BAheAAAoJEH4levi+sCboeHoH/3IyNGGwxVWy
VVnjKj2vpbgysU4W4xieL9sWvMBFnKDhpHZsazBEhXnmhEbDouixZaFeMmul8C7J
2/5Ljync/fkPCKtyF+Ibovs3ALuHnY4Iu8vukxMbr7cmB1lOkVGxHIKcjGX4H9F7
6qnGYmJWpz+pgYIbq5xO07aCcwE9/EUQwh0MdDml0euRiDWio1HOM7XTVJJ7AmyX
MKroqF+Ik/93mSl4vGlKKqDhPr3hcxqFsE8LhHgMxeI2NGomhka064mwWqRpFf0f
ce33cWaFSgl+rRAqkQkZUdiMnbIj9P89OH/PqOQgaB/nXIVXmjMb6HluhJA/2ZnM
h+Y8e0pF2j+5AQ0EU8LJcQEIAK38D6Bnho1cennrFOVcCj1nmlG4UW9mWr2ox+WJ
QBEqw8IUsWg/0LEe5K3MPoOE3lO3VTHnKqLMJSbA9byjSwYxIE3Y1QoY1Uq43Da1
sYkeETVkMgnAGmIwSQgsdfdAGhXv5uF/Ck3O+QMdUW5qZ4s+WXUCMWcj4ZgomUxC
i0bQgE/w/TDc0JAisma2oOuOTVjpfyX5VCk6XtwmDxE+STHZTCIKvSnyodx3Hlke
1gr+f/ejpbAnYiyjjWpiQGS47YCjAzGAsn0CRJ4dQYsjv6RVL/O/EYEJDUs49cLW
ccYhj0BQZyeMqWxZqP8ZUIGsoPfzh/ahLNXnxlqVTD53eA8AEQEAAYkBHwQYAQIA
CQUCU8LJcQIbDAAKCRB+JXr4vrAm6OFPB/9P4GkEV+XpejL4TO17Sh7vj3nZvKxd
AoPKKG1qbJNuYqasz0d5C0hfZN4aLaKdiWide9sIMfjRrG1gbN8o34uR3i3887Eg
zrhZWS/E01jGqR4ey/iACyfXvDfEFEwthfChyS9qQVYw7fWWSBtpZqJ5iul7Jf7b
tHPeqizK2FqOSnJy9ovaHHcZL4Wt26Y+IDWq0WQKB89guhN6LhlaQQXrAhlbwW2N
DcvTrHm3g4sVxeuAujGJzJGRmf5hkV+YxG2OrpLQjx+n4XsZSFO3tdfNwTwDn1Xj
9AFqGhRzm9j1Cq3iqcTbJtQwwJknkNm7CLFeHuy4zurzP3gmwnRvZ2UM
=X1Q+
-----END PGP PUBLIC KEY BLOCK-----

To the untrained eye someone's cat has been traipsing about on their keyboard, but in reality this string represents two (big, random) numbers that allows you to communicate with it's provider in a completely secure fashion. It is a "public key" and is used to encrypt messages sent to its owner; it has a mathematically precise twin, a "private key" which only the owner has access to. This private key is required to decode the message originally encoded by the public key.

For an extremely helpful analogy, consider a padlock. A public key can be thought of as an open padlock and a private key is the key that can unlock it once it is closed. In public key cryptosystems, you give your close friends copies of your personal padlock, open and unlocked (this is your public key). They can now send you messages securely by locking them in a box using your padlock (i.e., encrypting it with your public key) and sending it to you, because only you have the key (the private key) which can open it. Note that your friends don't need a key to close the padlock: they simply put their message in the box and shut the padlock.

How does this work mathematically? Here's a technically correct but very basic run-through of RSA for didactic purposes:

Let's say you want to communicate privately with another party. They're going to send you a secret message (message: "PRIVATE") over the internet, and they want to be sure that no one else can read it. In order to do so, they can encrypt this message using your public key and send it to you; you can then use your private key to decrypt and read it.

Let's assume you haven't yet generated public/private keys, and you want to do so by hand:

STEP 1
Choose a pair of prime numbers, p and q, at random.
To keep the math reasonable, let's take p=3 and q=17

STEP 2
Multiply p and q together to get their product n.
Here n = 3 x 17 = 51 (This is the step that is hard to reverse in practice!)

STEP 3
Multiply (q-1) x (p-1) to get the totient of n, φ(n)
φ(n)= (3-1) x (17-1) = (2) x (16) = 32.

STEP 4
For the first key, choose any number e that is smaller than φ(n) and has no common factors with φ(n). Since φ(n) = 32, we can choose any number besides 1, 2, 4, 8, 16, and 32.
Lets pick e =11.

STEP 5
Finally, the matching key d must be computed. This is achieved by "taking the inverse of e modulus φ(n)". All this means is that we need to find the number d to multiply e = 11 by so that when we divide their product by φ(n) =32, we get a remainder of 1. This isn't as hard as it sounds: φ(n)=32, and e=11. Since 11 x 3 = 33, and 33 divided by 32 leaves a remainder of 1, we know that the d, "the inverse of 11 modulus 32" , equals 3.

These keys e=11 and d=3 are mathematically linked through n=51, because
if you take the number you want to encrypt to the power of e and divide by n, you get a remainder. This remainder the encrypted version of your original number. To decode it, just raise it to power of d and divide by n.

As quick example, say the secret message you want to send is the number 4. To encrypt it using our public key (e=11, n=51), do

411 % 51 = 4194304 % 51 = 13.

The number 13 is our encrypted message, which can only be unlocked if you have the private key (d=3, n=51). The same procedure used to encrypt is used to decrypt, but take 13 to the power of d and find the remainder when dividing by 51.

133 % 51 = 2197 % 51= 4, our original message

Without factoring n = 51  you can't easily compute φ(n) 32 and thus knowing one key, you can't easily compute the other.



But how do we send the secret message "PRIVATE"? First you should get a numerical representation of this message; commonly, a much longer messages is being sent and it is converted from ASCII to its decimal representation. For now, we can just take the number that corresponds to each letter's position in the alphabet, (A=1, B=2,..., Z=26). Doing so for this message yields "16 18 9 22 1 20 5".

To encrypt our message ("16 18 9 22 1 20 5") using one of the keys (now therefore the public key), we repeat the process used above for the number 4, but now we use it on each of the numbers in our numeric code:

encrypted = originalpublic_key mod n.

1611 mod 51 = 17,592,186,044,416   % 51 = 16
1811 mod 51 = 64,268,410,079,232   % 51 = 18
911   mod 51 = 31,381,059,609         % 51 = 15
2211 mod 51 = 584,318,301,411,328 % 51 = 28
111   mod 51 = 1                               % 51 = 1
2011 mod 51 = 204,800,000,000,000 % 51 = 41
511   mod 51 = 48,828,125                % 51 = 11

So our encrypted message is (16 18 15 28 1 41 11). This is the message we send to our intended recipient. Even if it is intercepted in transit, it remains unintelligible without the private key.

To decrypt the message, repeat the process except now we are raising the encrypted message to the power of the private key (3), which transforms it back into its original code.

original = encrypted(private_key) mod n.

163 mod 51 = 4096   % 51 = 16
183 mod 51 = 5832   % 51 = 18
153 mod 51 = 3375   % 51 = 9
283 mod 51 = 21952 % 51 = 22
13   mod 51 = 1        % 51 = 1
413 mod 51 = 68921 % 51 = 20
113 mod 51 = 1331   % 51 = 5

Resulting in the original message, (16 18 15 28 1 41 11 = "P R I V A T E").


The way this works in practice is that you generate your own set of keys, one public and one private. The public key is made known to others with whom you wish to communicate privately (often by posting it somewhere online). Then, if someone wants to send you an encrypted message, they simply encode their message using your public key and send it to you. At this point, the message is garbled and can only be decoded using your private key. Remember that big block of garbled nonsense above? That's my public key, analogous to (11, 51) in the example above except that in decimal form it has over 300 digits!

In the demonstrative example above, our n = 51. Numbers like 51, which are the product of two primes numbers, are called semiprimes. Its not hard to see that 51 = 17 x 3, and these factors are all you need to crack our code! So how is this secure? The strength of the security offered by RSA and similar cryptographic methods is that finding the original factors of a huge semiprime is computationally difficult. For small semiprimes its no big deal, but when the two prime factors are large (~300 digits, which is more than a "googol"!), randomly chosen, and about the same size, the search becomes impractical for even the most powerful computers. The number of operations required to perform the factorization exhausts all our of present computer power.

The largest RSA number that has even been successfully factored is 768 bits (232 decimal digits), and this took hundreds of computers more than two years to accomplish! Indeed, much smaller RSA numbers, many with large bounties in their day, remain unfactored. Still, this method of encryption is not "uncrackable" and the size of the numbers used will have to stay one step ahead of developments in computing power. There do exist uncrackable codes, however...

Now, why are Fermat primes (22n+1) useful in RSA cryptography? Often, the public key exponent is one of these five numbers, typically 65537. Consider their binary representation:

3 = (11)2
5 = (101)2
17 = (10001)2
257 = (100000001)2
65537 = (100000000000000001)2

They are computationally convenient! There are probably other reasons too... let me know if you think of any!