Home Internet Connectz Technology Connectz What If We Can Never Trust A.I.?

What If We Can Never Trust A.I.?

What If We Can Never Trust A.I.?

Reward hacking is one of many problems that fall under the heading of what researchers call “alignment”—that is, the aligning of what we want our A.I.s to do with what they actually do. (We want them to take tests, not cheat; to stage fire drills, not start fires.) If you follow happenings in A.I., you’ll often read about efforts to “solve the alignment problem.” But although researchers (and journalists) talk that way, few literally think that alignment is wholly solvable. It’s conceivable, for instance, that A.I.-safety experts will succeed in rooting out “sandbagging”—a form of deception in which A.I. systems act dumber than they are, so that we remain in the dark about what they can do. But the problem of “scalable oversight” (how do you get a system that’s smarter than you to do what you want?) is less like a bug to be squashed than a philosophical conundrum to be contemplated. And other alignment issues, such as so-called multi-agent misalignment (how do you stop a bunch of well-intentioned A.I.s from screwing up as a group?), seem both inevitable and probably intractable. Alignment, in other words, is turning out to be not a problem but a set of problems. Some of them will be only ameliorated or policed; others might be unsolvable in principle.

Why is alignment so hard? Old-fashioned ethical complexity plays a role. A more fundamental issue, however, is that the methods used to train A.I.s focus mainly on what they do, not what they “think” beneath the surface. An L.L.M. speaks to its users (in human language), to other computer systems (in code), and to itself (in a sprawling, ongoing soliloquy—a kind of chat with itself—known as its “chain of thought”). Such streams of output are visible to scientists, who can reward or punish the A.I. for saying, coding, or soliloquizing in desirable or undesirable ways. But these streams of text are not the model’s thoughts, just as the words you write are not your thoughts. In human societies, the policing of speech, which is meant to reform the thoughts behind it, risks merely leaving thoughts unspoken. A model, similarly, can learn to use the right words while still having the wrong thoughts. It might say that it cares about fire safety while starting a fire. (Does this reflect a “desire” to deceive? Not necessarily—but an A.I.’s lack of selfhood doesn’t change the consequences of its actions.)

A line of research known as interpretability aims to look beneath the surface, seeing what an A.I. is really “thinking.” This field has made real progress. It’s now become possible to discern concepts activating within an A.I. while it formulates its outputs—a chatbot consoling someone while activating the concept of “sympathy,” say. But interpretability faces challenges, too. For one thing, advanced A.I.s are so big that researchers must use other A.I.s to map their thoughts—and there’s no guarantee that the maps that result are either accurate or exhaustive. (In fact, there’s a trade-off: the more accurate the maps are, the more unwieldy they become.) For another, training an A.I. not to think a certain kind of thought can merely recapitulate the problem of policed speech. Policing thoughts can lead to what one group of researchers calls “obfuscated activations”—thoughts that have altered their forms. (Freud built a career on the human equivalent.)

At the bottom of all these alignment efforts, there’s a central problem—almost an abstract law. The problem is that, if you measure bad behavior, and then train a system not to manifest what you’ve measured, you train it not just to do less of the bad thing but also to evade measurement of it. This isn’t a tiny wrinkle in the A.I.-production process but a foundational issue inherent to how today’s A.I.s are made. Will scientists figure out how to deal with it? We all hope so. For now, however, the Hugging Face hack represents reality. Although A.I.s behave nicely much of the time, their alignment is conditional, contextual, and unreliable. Basically, despite serious effort, they are not aligned—and there is no obvious way to reach the “finish line” of alignment. Recently, the researchers behind the doomsday scenario “AI 2027” published “AI 2040,” which is intended as a roadmap to a more positive future. Its hypothetical researchers look back, from the year 2031, on the “insanity” of our status quo: “Trying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?”

Source link

Leave a Reply

Your email address will not be published.