Undergraduate Research Assistant
Institute of Computing - University of Campinas
A two-stage sequential decision-making framework combining offline reinforcement learning, contextual bandits, and multi-objective optimization to eliminate harmful feedback loops, mitigate selective labels (reject inference), and ensure long-term fairness and policy flexibility.