Neuroscience and machine learning

Fly brain · hotdog or not hotdog

Research in progress
On this page

The question

What happens when a learning agent’s recurrent wiring is constrained by a fly-brain connectome?

What I’m exploring

This experiment uses the FlyWire fly-brain connectome as the structure for a recurrent neural network. Sensory inputs and action outputs connect to that network, which learns to navigate a food arena and distinguish hotdogs from other objects.

Sensory inputs pass through a connectome-constrained recurrent network to arena actions

Learning and interaction

The approach combines a leaky recurrent network with recurrent PPO reinforcement learning. Exported policy weights can run in a browser-based 3D arena, making the experiment something to watch and interact with.

Research direction

It explores the relationship between biological connectivity and learned behaviour. The network is a simplified computational experiment, rather than a reconstruction of everything a living fly’s brain does.

From connectivity to a learning system

The FlyWire FAFB v783 connectome provides a graph of neurons and their connections. The experiment maps sensory inputs and action outputs onto a selected network structure, then updates a simplified recurrent state as the agent observes the arena.

A leaky update preserves some previous activity while introducing the next step. This creates a memory-bearing policy rather than treating each observation as unrelated. Recurrent PPO trains the policy through sequences of interaction, with the arena providing observations, actions and rewards.

What the arena makes visible

The food-recognition task gives the model something concrete to do: move, encounter objects and make a decision about them. A browser-based 3D presentation makes that loop easier to inspect. Exported weights run locally in the visitor’s browser, connecting a trained policy to visible behaviour.

The visual simulation and the policy have separate roles. A convincing animation does not by itself demonstrate what the network learned. The useful evidence comes from repeated episodes and comparisons under consistent conditions.

Questions I want to investigate

Does the connectivity constraint help the task, hinder it, or mainly change how the policy learns? How much of the outcome depends on the selected input/output ports, recurrent dynamics or reward design? Would a simpler recurrent baseline perform similarly at the same budget?

Those comparisons matter because the graph supplies an inductive bias, not a complete biological model. The experiment omits much of real neuronal physiology, sensory processing and a living animal’s environment.

Connection to my other research

This project shares an interest in constrained computation with Ostinato, while using a different learning objective. It also connects neuroscience, graph representations and interactive software: the model’s behaviour can be studied through a small application rather than only through a training curve.