Reinforcement learning improves an agent through trial and error: it acts in an environment, receives rewards or penalties, and adjusts its strategy to maximise the gain. In the form of RLHF (reinforcement learning from human feedback), it is also a key step in aligning large language models.
In practice at Gensai
Gensai follows these techniques closely: they explain much of the behaviour of the models the agency integrates.