Abstract
Most systems that adapt to a user learn from labels: explicit ratings, reward signals, or annotated targets. We describe and evaluate a mechanism that learns from behavior alone, with no labels of any kind. The system acts, observes what the user does next, and updates its model in proportion to how surprising that behavior was — treating the user's own subsequent behavior as the only training signal. We call this reflection-on-action (e+r): act, then learn from the behavioral outcome. We formalize the label-free update (behavioral surprise as the sole learning signal, with the target always the externally produced behavior so the loop cannot reinforce its own hallucinations), state falsifiable hypotheses in advance, and run a pre-registered, held-out, non-circular experiment on 22,201 need-sequences. Results, reported verbatim including a refuted prediction: label-free behavioral learning beats a context-free frequency prior on next-need prediction (MRR 0.138 vs 0.052, ~2.6x, supported); the error-weighting term helps marginally; hop-1 propagation was predicted to help but was refuted (it slightly hurt); and the effect size triggered a pre-committed re-audit that narrows the claim to label-free conditional-from-behavior learning rather than person-specific learning. We argue the mechanism's central property — it is objective-agnostic — is also its central danger: pointed at engagement rather than wellbeing, the same learner maximizes dependency, and does so most effectively on the users least able to afford it. We treat that as a result, not a misuse footnote, and specify the wellbeing-gated deployment it requires.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Sanyal, Arko, "Reflection-on-Action: Label-Free Learning from Behavioral Outcomes in Interactive Agents (LOCI e+r)", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11090