In search of better models for learning in perceptual decision making: How stimuli and rewards shape behaviour

Author: Christina Koß

Referees:
Prof. Dr. Frank Jäkel
Prof. Constantin Rothkopf, Ph.D.

Defense: 11.08.26

Abstract:

Perceptual decision making (PDM) deals with situations where humans or other animals have to decide on an action based on some sensory input as well as expected positive and negative outcomes. In a typical lab setup to study this behaviour, subjects complete a categorisation task in a series of trials. During each trial, a stimulus is presented and the subject has to indictate which of two categories it belongs to. Correct responses are intermittently rewarded. Categories overlap, so the subject cannot know with 100% certainty which response is the correct one. In order to perform well and collect as many rewards as possible, subjects adapt their behaviour to the situation, which is especially beneficial in dynamic environments where stimulus or reward frequencies change. The mechanisms behind their learning process are not fully understood yet.

Our goal is finding an algorithmic model that describes a potential learning mechanism and can generate the same behaviour as the animals. We model PDM using signal detection theory (SDT), which posits that subjects decide by comparing the perceived evidence to a decision criterion. To model the learning process, we use criterion learning models, i.e., a class of SDT models where the position of the criterion can shift from trial to trial.

An investigation of three existing criterion learning models that integrate past rewards (IR model), reward omissions (IRO model) or both (IR&RO model) revealed that they could not adequately describe the animals’ behaviour in two experiments with varying reward probabilities.

The data from these two experiments showed a relation between the amount of reward received for each category and the animals’ response rates, following regularities that are often observed in two-stimulus setups and have been formalised by Davison and Tustin (1978) in what we call the Davison-Tustin (DT) law.

Therefore, we developed a model that is consistent with the DT law – the DT model. We derived the criterion position that would produce behaviour according to the law, and designed a rule for trial-by-trial criterion updates that leads to this position in the steady state.

To model behaviour in a multi-stimulus setup, where the DT model is not appropriate, we modified an existing criterion learning model. We conducted an experiment designed to identify which outcomes cause the animals to shift their criterion and found that animals learned from rewards rather than reward omissions. Moreover, we found that steady-state behaviour was independent of overall reward density, and that the change in animals’ response bias was larger following more difficult trials. We modified the IR model such that it captured these phenomena.

We evaluated the DT model and the modified KDB model, as well as another model based on ideas from reinforcement learning (RL model) for comparison, on a series of experiments that systematically investigated the effects of stimulus presentation probabilities (SPPs) and reward probabilities (RPs). The DT model and modified KDB model could explain data well when either SPPs or RPs were varied, but failed to provide a consistent account for joint variations of SPP and RP. The RL model included fixed assumptions about the stimuli and therefore could only explain the data in the experiment with constant SPPs.

In summary, incorporating empirical findings that were unaccounted for by previously existing criterion models allowed us to develop models that explain behaviour well in a range of experimental setups. However, these models cannot capture the differential effects of stimuli and rewards on animal behaviour. We conclude that models for learning in perceptual decision making should represent stimulus properties and reward expectations separately and adapt both representations based on the experienced percepts and response outcomes, instead of using fixed stimulus distributions and representing the whole learning process as a decision criterion that is updated.