Research

We evaluate theories about agency according to their scientific merits — the naturalistic plausibility of their assumptions, their conduciveness to generating falsifiable predictive models, and the computational feasibility of those models. Most importantly, we run experiments to test the key hypotheses implied by the most promising theories.

Theories we are currently evaluating

  • Expected Utility Theory (EUT)
    • The von Neumann–Morgenstern representation theorem defines plausible conditions (axioms) under which goal-directed behaviour can be described as maximising expected utility.
    • Violating vNM's axioms can create some vulnerability to guaranteed losses (money pumps), which suggests pressure to satisfy the axioms.
    • Goodharting frequently appears as a major risk in EUT.
    • This involves maximising utility under a definition which imperfectly captures what is actually desired (a proxy), due to the difficulty of specifying the extremely complex and often underspecified preferences of realistic agents like humans.
    • EUT is often defended on a normative basis, as a standard of behaviour we should aspire to.
  • Infra-Bayesianism (IB)
    • Infra-Bayesianism is a generalisation of both Bayesianism (for beliefs) and Expected Utility Theory (for behaviour) to cases where there is Knightian uncertainty.
    • This occurs when there is no reasonable single choice of a prior probability distribution, so a range of (unweighted) possibilities must be considered simultaneously.
    • In order to avoid worst-case risks, Infra-Bayesianism assigns scores to outcomes in a way that is concave with respect to the probabilities.
    • As such, they violate the 4th vNM axiom (independence), but research suggests that they are still invulnerable to money pumps (due to a form of updatelessness or resolute choice).
    • There are methods of aggregating group preferences for social choices which (in general) can only be represented using Infra-Bayesianism, and not EUT, such as Nash bargaining.
  • Active Inference (AIF)
    • Active Inference treats both beliefs and predictions as the same type of mathematical object - a probability distribution.
    • From this perspective, prediction errors can be corrected either by modifying beliefs or by modifying the world, depending on the confidence in the prior prediction.
    • A number of other ingredients also influence behaviour, including self-models (particularly persistent beliefs about identity) and information-seeking.
    • Boundaries between individuals are identified not based on behavioural algorithms but on Markov blankets.
    • A Markov blanket is defined such that the inside of the blanket (internal states) doesn't provide any additional value beyond the blanket itself for predicting the outside (the environment).

Select a theory to read more.

Theories we are interested in exploring

We would love to hear from researchers knowledgeable about any of the following.

auer.ben@proton.me

Outputs

Nothing published yet — this section will list papers, preprints and talks as they appear.