The rat had learned the maze so thoroughly that a tone could send it left or right without the value of the destination seeming to matter anymore. Then yellow light reached a tiny patch near the bottom of its medial prefrontal cortex. Within seconds, the old routine stopped dictating the turn.
This was an optogenetics experiment in ten rats, published in 2012, not a human study and not a technique for switching off an unwanted behavior in everyday life. Its value was narrower and more revealing. A well-established habit could lose control almost immediately when one cortical region was interrupted, yet the habit itself remained available and later returned.
In other words, stopping a stored behavior from running is not the same as deleting it. The paper in the Proceedings of the National Academy of Sciences, led by MIT neuroscientist Kyle Smith, gave that distinction a causal test measured in seconds.
First, the researchers had to show that the rats were acting from habit
Smith, Arti Virkud, Stanford bioengineer Karl Deisseroth and MIT Institute Professor Ann Graybiel trained rats on a T-shaped maze. At the start of each run, a warning cue told the animal to begin. A second auditory cue instructed it to turn towards one of two goal arms. A correct turn led to about 0.3 millilitres of chocolate milk or sucrose solution, with the cue and reward locations fixed for each rat.
The ten animals completed roughly 40 trials in a daily session. They first had to reach a performance criterion of 72.5 percent correct, then continued for at least ten overtraining sessions. Peak accuracy was about 90 percent.
Accuracy alone does not establish a habit. A rat could keep making the correct turn because it still wanted the outcome and was deliberately choosing the route that delivered it. The team therefore used reward devaluation, a standard way to separate goal-directed performance from behavior that has become relatively insensitive to its result.
One reward was offered in the home cage and paired three times with lithium chloride, which produced an aversion. Later, the rats were returned to the maze for an unrewarded probe. If they were choosing according to current value, they should avoid the arm associated with the now-unpleasant reward. If the cue-to-turn routine had become habitual, they should continue following the instruction despite no longer valuing what once waited at the end.
Four control rats kept running towards the devalued goal when cued, even though separate consumption tests showed that they now avoided that reward. That insensitivity to the changed outcome was the operational evidence that overtraining had produced a maze habit.
Yellow light interrupted the habit during a single run
The critical animals had been given a viral vector that made pyramidal neurons in the infralimbic cortex express eNpHR3.0, a version of the light-sensitive chloride pump halorhodopsin. Optical fibres were implanted above the region. Shining yellow light at 2.5 to 5 milliwatts suppressed local population activity.
The light was not left on for hours or even minutes. It began with the warning cue and ended when the rat reached the goal, usually about three seconds later. This timing let the researchers ask whether the cortex was needed online, while the habitual action was actually unfolding.
The change arrived within an average of three trials, equivalent to roughly nine seconds of inhibition in total. In some cases it appeared during the first illuminated run, in under three seconds. Rats carrying the active light-sensitive protein reduced or largely stopped going to the devalued side. Instead, they often turned towards the still-valued goal even when the tone instructed the other direction.
The animals had not simply frozen or forgotten how to navigate. Runs to the non-devalued goal remained intact, and the researchers did not detect a general loss of performance ability or motivation to consume reward. The brief cortical disruption had changed which information controlled the choice.
MIT’s contemporaneous report on the study described this as a shift from an automatic mode towards behavior more engaged with the current goal. That is a useful interpretation, although the behavioral test cannot reveal anything like a rat’s conscious deliberation.
The replacement behavior became automatic too
After the first probe, the rats returned to training. Over the following weeks, animals in the inhibited group developed a different pattern. They increasingly ran towards and consumed the still-valued reward, even when the auditory instruction pointed towards the devalued side. What began as sensitivity to the new value structure became a stable replacement strategy.
The researchers then applied the same brief infralimbic inhibition again. This time the result ran in the opposite direction. The newer pattern was blocked and the original cue-guided behavior reappeared, including runs towards the devalued goal. The transition took an average of two trials, or about six seconds of total inhibition.
That return is the strongest evidence behind the title’s claim that the first habit remained hidden. If the initial intervention had erased the old cue-to-turn association, suppressing the replacement weeks later could not have restored it almost immediately. The earlier routine had ceased to govern behavior, but some usable representation of it had survived.
The paper also acknowledged an interpretive complication. The late change could be described as a return to a value-driven state rather than literal reinstatement of an old habit. Reward consumption and running patterns, however, led the authors to favor the reinstatement account. The experiment supports a hidden surviving strategy; it does not provide a direct recording of a memory being stored untouched.
The cortex looked more like an arbiter than an archive
Habits are commonly associated with the basal ganglia, especially circuits involving the sensorimotor striatum. Earlier work from Graybiel’s laboratory had found a characteristic “task-bracketing” pattern there: neuronal activity became strong near the beginning and end of a practiced sequence while quieting through its familiar middle.
The infralimbic cortex sits in a different position. Lesion and inactivation studies had already suggested that this medial prefrontal area was needed for habits to be expressed. The 2012 result added temporal precision. Its activity mattered during the few seconds in which the behavior was performed, and interrupting it could rapidly change which learned strategy won control.
Calling the region a switch is therefore a metaphor for function, not a claim that one clump of cells contains an entire habit. The team’s model placed durable action representations within a wider corticostriatal system. Infralimbic activity appeared to favor the currently dominant habitual strategy, particularly the more recently acquired one.
A follow-up study in Neuron added another piece. As rats acquired and lost habitual performance, activity patterns in the infralimbic cortex appeared and disappeared with the habit state. Inhibiting the region during overtraining could prevent a new habit from crystallizing, even while the rats still learned the maze task. Selection and formation were intertwined, rather than the cortex merely issuing a final command to a fully independent routine.
There is no human habit-off button in this experiment
The study’s control was impressive because it was invasive and unusually specific. The animals received a viral construct, surgically implanted optical fibres and precisely timed light. Nothing in the experiment resembles a non-invasive action a person could take when trying to change a routine.
There is also no clean one-to-one translation from the rat infralimbic cortex to a single human brain region with the same boundaries and role. Human habits are shaped by language, expectation, social setting and years of learning. Smoking, compulsive behavior and substance use are not larger versions of choosing an arm in a T-maze.
The small sample matters as well. Ten rats were enough to identify a causal effect under tightly controlled conditions, but not enough to establish that every habit or every animal uses the circuit identically. Later research has refined the picture of infralimbic function, including evidence that it may help maintain the currently active behavioral state rather than serving as a universal habit controller.
This is one rat experiment, not settled consensus about human behavior. It says that a particular cortical node could govern expression of a learned maze strategy second by second. It does not say that suppressing the analogous tissue in a person would safely remove a habit, nor that an unwanted human routine remains stored in precisely the same form.
What “old habits die hard” meant in this maze
Automatic behavior is often described as though control has disappeared. The rats suggest something subtler. Once a sequence became efficient, its detailed execution could run with little flexibility, while a higher-level circuit still helped decide which sequence occupied the driver’s seat.
That arrangement is useful. A brain can preserve well-practised behavior without rebuilding it on every occasion, yet retain some capacity to switch strategies when value or context changes. It also creates persistence. Suppressing one behavior does not require dismantling every neural trace that could later produce it.
The 2012 experiment made that persistence visible by accident of sequence. First, light unmasked flexible choice beneath an old habit. Weeks later, the same light unmasked the old habit beneath its replacement. The behavior that vanished had not necessarily been forgotten. It had been waiting outside control.