BNY AI Lab @ CMU: First Cohort of Awardees

10 Awardees Announced for 2026

Home > Research > BNY AI Lab @ CMU: First Cohort of Awardees

The BNY AI Lab @ CMU invited Carnegie Mellon University faculty to submit proposals for one-year research projects supporting PhD students or postdoctoral researchers aligned with the Lab’s mission of advancing Reliable, Responsible, and Resilient (RRR) Agentic AI for mission-critical systems.

Proposals were solicited to address the scientific and engineering foundations of trustworthy autonomous AI, particularly in areas including:

  • Autonomous agent verification and constraint enforcement
  • Evaluation science and stress testing for agentic systems
  • Multi-agent coordination, incentive alignment, and mechanism design
  • System-level robustness and failure cascade modeling
  • Continual and adaptive learning under governance constraints
  • Responsible AI by design (auditability, traceability, calibrated uncertainty, human-in-the-loop oversight)
  • Secure deployment, adversarial resilience, and model update integrity

Projects were encouraged to engage with the current state of the art in large models and agentic systems and to address challenges that arise in real-world, mission-critical environments.

Interdisciplinary collaborations across machine learning, systems, robotics, security, economics, public policy, and human-computer interaction are particularly encouraged.

Sarah Cen, Assistant Professor and PI for one of the ten projects that were awarded, noted “These funds will go towards a project that my students and I are really excited about. This project could jumpstart a new direction on AI agent evaluation and without these funds, we wouldn't have the flexibility to explore this new direction.”


The ten projects to receive this first award from the BNY AI Lab @ CMU are:

Deep Research Agents for Open-Ended Financial Research in Mission-Critical Settings

PI: Akari Asai, Assistant Professor, Language Technologies Institute, School of Computer Science

Project Summary: Researchers will develop next-generation AI research agents that can autonomously analyze complex financial information while adapting to new data sources, tools, and changing environments. By creating new training methods that improve how these agents gather evidence, reason through open-ended problems, and explain their decision-making, the project aims to make AI assistants more reliable, transparent, and resilient for mission-critical financial applications.

Bidirectional User-Agent Modeling for Reliable and Predictable Human-Agent Interaction

PI: Jeffrey P. Bigham

Project Summary: Rather than simply improving task completion, this project focuses on helping AI agents and people better understand one another. By developing AI systems that can model a user's goals and expertise, and helping users better understand how AI agents make decisions, the research aims to foster more predictable, trustworthy, and effective human-AI collaboration in complex, essential tasks.

Developing an Evaluation Science: Predictive Models for AI Agent Evaluation

PI: Sarah H. Cen, Assistant Professor, Departments of Electrical & Computer Engineering and Engineering & Public Policy

Project Summary: Instead of relying on costly, application-specific testing, this project explores whether AI agents' behavior can be predicted before they are deployed in new environments. The resulting framework will help organizations anticipate failures, target evaluations more efficiently, and improve the reliability of AI systems used in crucial settings.

“An Army of Me”: AI-Powered Digital Twins for Replicating and Scaling Human Work Practices

PI: Motahhare Eslami, Associate Professor, Human Computer Interaction Institute, School of Computer Science; Maarten Sap, Language Technologies Institute, School of Computer Science

Project Summary: Researchers will create AI-powered digital twins that learn how people reason, communicate, and make decisions so they can serve as personalized assistants for complex professional tasks. By modeling expert judgment, not just task completion, and by incorporating human oversight, the research aims to build AI agents that provide reliable, transparent, and context-aware decision support.

Toward Valid GenAI Evaluations: An Efficient Hybrid Approach for Operationalizing Context-Specific Dataset Generation

PI: Hoda Heidari: Assistant Professor, Machine Learning Department, School of Computer Science

Project Summary: To improve confidence in AI systems, this project will develop scalable methods for building benchmark datasets tailored to specific real-world applications. The approach combines synthetic data generation with strategic expert review to create more reliable evaluations while reducing the time and effort required from human experts.

Scalable Oversight of AI Agents via Trustworthy Reasoning Traces

PI: Daphne Ippolito, Assistant Professor, Language Technologies Institute, School of Computer Science

Project Summary: To improve trust in AI systems, this project will explore how to make the reasoning behind AI decisions more transparent and easier to verify. The resulting methods will help identify errors, hallucinations, and deceptive reasoning, supporting safer deployment of AI agents.

Grounded Agentic Translation of Legacy Systems

PI: Claire Le Goues, Software and Societal Systems Department, School of Computer Science

Project Summary: This project seeks to modernize legacy COBOL software by combining AI with program analysis to translate decades-old code into reliable, maintainable Java applications. By grounding AI-generated translations in verified program behavior and data relationships, the research aims to make migration of financial systems more accurate, auditable, and resilient.

Formally Verified Efficient Code Generation through Agentic AI

PI: Elaine Shi, Computer Science Department, School of Computer Science

Project Summary: This project aims to make AI-generated code more trustworthy by combining AI with formal verification techniques that can mathematically prove code is correct. By enabling AI systems to automatically generate, test, and refine software against rigorous specifications, the research seeks to reduce human effort while producing reliable, efficient code for pivotal applications such as cryptography and distributed systems.

Detecting and Evaluating Multi-Agent Collusion

PI: Virginia Smith, Associate Professor, Machine Learning Department, School of Computer Science

Project Summary: COLLUSIONBENCH is intended to be an open-source benchmarking platform to detect and evaluate hidden collusion among AI agents operating in multi-agent systems. By combining advanced interpretability techniques with realistic testing environments, the project aims to uncover when AI agents coordinate in unintended or deceptive ways, even when that behavior is not visible in their outputs. The resulting tools will help improve the reliability, transparency, and safety of AI systems deployed in mission-critical domains such as finance, content moderation, and secure data access.

Incentive-Compatible Defenses for Trustworthy Multi-Agent AI Systems

PI: Chenyan Xiong, Associate Professor Language Technologies Institute, School of Computer Science

Project Summary: This project will strengthen the reliability of multi-agent AI systems by developing new ways to prevent and defend against untrustworthy behavior among collaborating AI agents. Combining adaptive defenses with incentive-based mechanisms that encourage cooperation, the research aims to ensure AI agents remain reliable, resilient, and aligned even in complex, high-stakes environments.