Skip to main content
NC State Home
Awards and Honors

Aritra Mitra Wins NSF CAREER Award

man with short dark hair wearing glasses and a dark, long-sleeve shirt; standing in front of two white dry erase boards

Congratulations to Department of Electrical and Computer Engineering Assistant Professor Aritra Mitra on receiving a Faculty Early Career Development (CAREER) award from the National Science Foundation (NSF). This prestigious award supports early-career faculty who serve as role models in both research and education. Mitra’s grant will support his project, “Foundations of Robust and Sample-Efficient Data-Driven Decision Making and Control.” 

“I am honored to receive this prestigious award from the NSF, which recognizes the value and potential impact of the work being performed by my research group,” said Mitra. “This award will allow our group to keep pushing the frontiers of robust reinforcement learning, contributing not only to new algorithmic methods and fundamental theory, but also to applications spanning wireless networks, robotics and power systems.”

Reinforcement Learning

Modern life and technology has advanced, and will continue to change, at a rapid pace. In particular, large-scale systems that involve groups of agents that interact with one another, like a fleet of drones or connected self-driving vehicles, are incredibly complex. 

Classical model-based control (where mathematical models are created of a physical system in order to predict behavior and optimize control) as a tool is not able to keep up with the complexity of modern machines. This has led to a rise in data-driven decision making paradigms known as reinforcement learning (RL), where an artificial intelligence agent is able to learn via trial and error based on feedback data. 

Essentially, reinforcement learning is the AI version of training your pet using positive reinforcement, i.e., treats, in order to achieve an optimal outcome.

RL has great upsides, but it also has a critical weakness: the data used to train a model must be correct and trustworthy. 

“Over the last decade, reinforcement learning has found tremendous use in a variety of applications, spanning finance, medicine, recommendation systems, autonomous driving, robotics — and most recently — training large language models (LLMs) using human feedback,” said Mitra. “Despite this progress, existing RL frameworks work under the idealistic assumption of perfect feedback data for training policies.” 

Perfect data does not exist in the real world. RL datasets can contain extreme noise (random spikes in data), out of distribution samples (situations an AI has not seen before), human bias and can be intentionally corrupted by adversaries (hackers or other bad actors trying to compromise an AI’s learning). 

Robust Reinforcement Learning

To address this critical shortcoming in reinforcement learning, Mitra and his research group are working to develop scalable and efficient decision-making algorithms that can function in the real world and come with provable guarantees on performance. 

To do this, Mitra and his team are combining reinforcement learning with robust statistics (where bad or fake data is ignored) to create AI systems that filter out corrupted information and make guaranteed safe choices. The CAREER project is broken into three thrusts:

  1. False Reward Filtering: Reinforcement learning is based off of a point or reward-based training system, where an AI agent is rewarded with a ‘treat’ for achieving a desired outcome. Attackers or bad data can corrupt the reward system, so Mitra is working to find the mathematical limit of how much bad data an AI agent can overcome. This also involves developing new, robust algorithms that will still make smart decisions when faced with noisy and potentially corrupted rewards.
  2. Teamwork and Communication: When multiple AI agents work together, communication costs create a key information bottleneck.. Mitra is working to ensure that a group of AI agents can still collaborate and safely share data under such bottlenecks, even if some of the AI agents are corrupted.
  3. Control in Changing Environments: Each large-scale system is unique and needs to function within a complex, constantly changing environment. Mitra is working to create data-driven control algorithms that work in unknown dynamic systems, so that if the real-world environment changes, the AI system can remain stable and still achieve the optimal outcome. 

Using autonomous vehicles as an example, Mitra’s work in this CAREER award can be used to:

  • Enable a fleet of self-driving cars to overcome bad data, hackers or flaws in a system.
  • Stay in communication with one another and safely share data about positioning, speed, movement and other conditions, even if some of the AI agents in the fleet are corrupted.
  • Work within a complex, ever changing environment. Roads close, people jaywalk, a thunderstorm comes in out of nowhere — a fleet of vehicles needs to be able to react quickly and prioritize human safety on any given street, anywhere in the world. 

Mitra’s research in this project has the potential to advance reinforcement learning in a variety of industries, from finance to healthcare and more. 

“Our modern world relies on large-scale systems operating autonomously in unpredictable environments,” said Mitra. “I am passionate about this research since it builds the mathematical foundations needed for such systems to process information from massive datasets to achieve robust real-time decision-making under uncertainty.”

This post was originally published in the Department of Electrical and Computer Engineering.