Skip to main content

liacc.com

Intelligent Robotics​ 

In the area of Intelligent Robotics LIACC has been researching in continuous black-box optimization and developed a battery of new, state-of-the-art algorithms to tackle continuous black-box optimization problems in the area of Robotics such as Trust Region Covariance Matrix Evolution Strategy (TR-CMA-ES) and Model-Based Relative Entropy Stochastic Search (MORE). The work was also applied in the FC Portugal robotic soccer team for the optimization of low-level and mid-level robot skills and helped FC Portugal team to achieve 10 international awards including the winner of three RoboCup scientific/free challenges. The ability for a robot to coordinate with others within a system is a valuable property in multi-robot systems.

Robots either cooperate as a team to accomplish a common goal or adapt to opponents to complete different goals without being exploited. Research has shown that learning multi-agent/multi-robot coordination is significantly more complex than learning policies in single agent/robot environments and requires a variety of techniques to deal with the properties of a system where agents learn concurrently. In this context we have been researching how can machine learning be used to achieve coordination within a multi-agent/multirobot system and what techniques can be used to tackle the increased complexity of such systems and their credit assignment challenges, how to achieve coordination, and how to use communication to improve the behaviour of a team.

Many algorithms for competitive environments are tabular-based, preventing their use with high-dimension or continuous state-spaces, and may be biased against specific equilibrium strategies. We have proposed multiple deep learning extensions for competitive environments, allowing algorithms to reach equilibrium strategies in complex and partially observable environments, relying only on local information. We also have developed multi-agent algorithms where agents learn communication protocols to compensate for local partial-observability and remain independently executed. A centralized learning phase can incorporate additional environment information to increase the robustness and speed with which a team converges to successful policies.