Science of Agency Research Institute
We are an independent, non-profit research organization working to accelerate the development of a new scientific paradigm in which falsifiable theories of agency can be developed, which make empirical claims about real, bounded agents. We approach this from an inter-disciplinary perspective, drawing on insights from cognitive science, economics, learning theory, and the history and philosophy of science.
Why agency?
We define agency to mean autonomous, goal-directed behaviour, and related phenomena. We believe understanding such phenomena is essential for addressing the AI alignment problem — the problem of ensuring that powerful, autonomous AI pursues goals that are compatible with human (and animal) flourishing.
With AI capabilities rapidly progressing, now is a crucial moment to develop a scientifically-grounded theory of which types of agents are robustly safe. In particular, we seek a thermodynamics of agency, a theory of macroscopic principles which do not appeal to specific details of the architectures of neural networks or brains.
Why science?
Science is our best tool for understanding the natural world, and the phenomena that occur within it, including goal-directed behaviour. We believe the hardest parts of AI alignment concern empirical questions, like what types of agents or goals generalise most reliably, and where mesa-optimisers arise. While mathematics and philosophy can help us formulate hypothetical answers to these questions, or systematically organise what we already know, the only way to truly push our empirical knowledge forwards is by running experiments which force our hypotheses to collide with reality. Only very few will survive the scientific process.
Our research
We evaluate theories about agency according to their scientific merits — the naturalistic plausibility of their assumptions, their conduciveness to generating falsifiable predictive models, and the computational feasibility of those models. Most importantly, we run experiments to evaluate the key predictions generated by those models.
Currently evaluating
- Expected Utility Theory (EUT)
- Infra-Bayesianism (IB)
- Active Inference (AIF)
Also interested in
- Viability Theory
- Live Theory
- Cybernetics
- Evolutionary Game Theory
Get involved
Working with us
If you're researching AI safety or anything related to agency, especially the theories above, and this framing resonates with you, we'd love to hear from you. We're also very open to constructive criticism. Email us.