**Background**
Appetitive associative learning—the process by which animals learn to associate cues with rewards—is fundamental to adaptive behavior, but its dysregulation can contribute to maladaptive states such as addiction. While much research has focused on aversive learning, the neural mechanisms underlying reward-based learning remain less understood. The medial prefrontal cortex (mPFC), particularly the prelimbic (PL) cortex, is implicated in both promoting and inhibiting reward-seeking behavior. This study aimed to characterize the neural populations recruited across days of appetitive associative learning and to causally test the role of the PL in encoding and recall of a reward-associated cue.
**Methods**
Sixty-two male C57BL/6 mice (2–3 months old) were trained in an operant conditioning task: a light cue (30 sec) signaled availability of 10% sucrose water via a nose poke, followed by a 60 sec light-off period with no reward. Mice underwent 3 days of water training (light always on) followed by 7 days of light training (15 alternating light-on/light-off trials). Behavior was quantified as time spent poking during light-on and light-off, and a discrimination difference (DD) was calculated: (poking during light-on/30) − (poking during light-off/60). Whole-brain cFos expression was assessed at four time points: after water training only (R0, no association), and on days 1 (R1, encoding), 2 (R2, recent recall), and 7 (R7, remote recall) of light training. A separate control group (Ctx) experienced the chamber and light without reward. cFos-positive cells were counted using the QUINT workflow. To tag engram cells, a doxycycline-dependent cFos-tTa/TRE-mCherry system was used to label PL neurons active during R1; their reactivation during R7 was quantified. Chemogenetic inhibition was performed using AAV5-hSyn-hM4Di-mCherry (or control mCherry) injected into the PL. Clozapine-N-oxide (CNO, 3 mg/kg) was administered 30 min before behavior to inhibit PL neurons either during R1 (early) or R8 (late, an additional day after R7). In a separate experiment, activity-dependent DREADD (AAV9-cFos-tTA-TRE-hM4Di-mCherry) was used to selectively inhibit only the R1-tagged engram cells during R8.
**Key Results**
- Mice learned to discriminate the cue: DD increased significantly across days (F(2,37)=6.627, P=0.0035, R²=0.2637). By day 2, mice poked more during light-on vs. light-off (t₇=2.536, P=0.039), and by day 7 the difference was more pronounced (t₆=3.363, P=0.0152).
- Whole-brain cFos analysis revealed distinct patterns: PL cortex showed peak activity at R1 and decreased by R7 (F(4,25)=6.437, P=0.0011, R²=0.5074). Other regions (e.g., hippocampus, caudoputamen, amygdala) also showed day-dependent changes.
- Linear regressions showed that PL cFos activity was negatively associated with DD at R1 but positively associated at R7, indicating that better performers at R7 had more PL cFos-positive cells.
- Engram tagging: 55.8% of PL cells active during R1 were reactivated during R7. The percentage of overlapping cells was significantly higher than chance (t₄=3.111, P=0.0358) and correlated with improvement in DD (R²=0.70, P=0.038).
- Chemogenetic inhibition of the entire PL during R1 had no effect on behavior at R1, R2, or R7. However, inhibition during R8 increased poking during light-off (t₈=2.600, P=0.0316) without affecting light-on poking.
- Selective inhibition of R1-tagged engram cells during R8 significantly decreased poking during light-on (t₈=2.753, P=0.0250) and reduced DD (t₈=2.664, P=0.0286), with no effect on light-off poking.
**Clinical Implications**
These findings demonstrate that the PL cortex plays a critical role in late-stage appetitive associative learning, particularly in fine-tuning cue-driven reward-seeking behavior. The dissociation between global PL inhibition (which increased inappropriate poking during light-off) and engram-specific inhibition (which decreased appropriate poking during light-on) suggests that distinct PL subpopulations mediate different aspects of reward learning. This has implications for disorders such as addiction, where maladaptive reward-seeking persists despite negative consequences. Targeting PL engram cells may offer a more precise therapeutic approach to modulate reward-related memories without broadly disrupting prefrontal function. Future studies should examine sex differences and explore other brain regions (e.g., habenula) that showed unique activity patterns.