{"owner":"Yizirui Fang","papers":[{"abstract":"Learning to defer (L2D) aims to optimize human-AI collaboration by allocating prediction tasks to either a machine learning model or a human expert, depending on which is most likely to be correct. This allocation decision is governed by a rejector: a meta-model that routes inputs based on estimated success probabilities. In practice, a poorly fit or otherwise misspecified rejector can jeopardize the entire L2D workflow due to its crucial role in allocating prediction tasks. In this work, we perform uncertainty quantification for the rejector. We use conformal prediction to allow the rejector to output prediction sets or intervals instead of just the binary outcome of ‘defer’ or not. On tasks ranging from image to hate speech classification, we demonstrate that the uncertainty in the rejector translates to safer decisions via two forms of selective prediction.","abstract_source":"https://openreview.net/pdf?id=SZQJ8K2DUe","abstract_source_label":"TMLR published article, February 2026, page 1","authors":["Yizirui Fang","Eric Nalisnick"],"bibtex":"https://yiziruifang.com/publication/fang-2026-learning-tmlr/cite.bib","pdf":"https://openreview.net/pdf?id=SZQJ8K2DUe","research_topics":["human-ai-decision-making"],"summary":"Conformal prediction quantifies uncertainty in learning-to-defer rejectors, enabling abstention and human-model consensus workflows for image and hate-speech classification.","title":"Learning to Defer with an Uncertain Rejector via Conformal Prediction","topics":["Learning to Defer","Conformal Prediction","Uncertainty Quantification","Human-AI Collaboration","Selective Prediction","Abstention"],"url":"https://yiziruifang.com/publication/fang-2026-learning-tmlr/","venue":"Transactions on Machine Learning Research (TMLR)","year":"2026"},{"abstract":"Spoken language instructions are ubiquitous in agent collaboration. However, in real-world human-robot collaboration, following human spoken instructions can be challenging due to various speaker and environmental factors, such as background noise or mispronunciation. When faced with noisy auditory inputs, humans can leverage the collaborative context in the embodied environment to interpret noisy spoken instructions and take pragmatic assistive actions. In this paper, we present a cognitively inspired neurosymbolic model, Spoken Instruction Following through Theory of Mind (SIFToM), which leverages a Vision-Language Model with model-based mental inference to enable robots to pragmatically follow human instructions under diverse speech conditions. We test SIFToM in both simulated environments (VirtualHome) and real-world human-robot collaborative settings with human evaluations. Results show that SIFToM can significantly improve the performance of a lightweight base VLM (Gemini 2.5 Flash), outperforming state-of-the-art VLMs (Gemini 2.5 Pro) and approaching human-level accuracy on challenging spoken instruction following tasks.","abstract_license":"CC BY 4.0","abstract_license_url":"https://creativecommons.org/licenses/by/4.0/","abstract_source":"https://arxiv.org/abs/2409.10849v2","abstract_source_label":"arXiv preprint 2409.10849v2, October 6, 2025","authors":["Lance Ying","Xinyi Li","Shivam Aarya","Yizirui Fang","Yifan Yin","Jason Xinyu Liu","Stefanie Tellex","Joshua B. Tenenbaum","Tianmin Shu"],"bibtex":"https://yiziruifang.com/publication/ying-2024-siftom/cite.bib","pdf":"https://yiziruifang.com/publication/ying-2024-siftom/paper.pdf","pdf_source":"https://arxiv.org/pdf/2409.10849v2","research_topics":["embodied-instruction-following"],"summary":"SIFToM combines vision-language models with probabilistic theory-of-mind inference for noisy spoken instruction following in human-robot collaboration.","title":"Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind","topics":["Spoken instruction following","Theory of mind","Human-robot collaboration","Neurosymbolic AI","Vision-language models","Pragmatic goal inference"],"url":"https://yiziruifang.com/publication/ying-2024-siftom/","venue":"IEEE International Conference on Robotics and Automation (ICRA 2026)","year":"2026"},{"abstract":"Integrating Large Language Models (LLMs) in educational technology reveals unprecedented opportunities to improve instructional design (ID), yet current approaches often prioritize automation over pedagogical rigor and human agency. This paper introduces ARCHED (AI for Responsible, Collaborative, Human-centered Education Instructional Design), a framework that implements a structured multi-stage workflow between educators and AI. Unlike existing tools that generate complete instructional materials autonomously, ARCHED cascades the development into distinct stages, from learning objective formulation to assessment design, each guided by Bloom’s taxonomy and enhanced by LLMs. This framework employs multiple specialized AI components that work in concert: one generating diverse pedagogical options, another evaluating their alignment with learning objectives while maintaining human educators as primary decision-makers. ARCHED addresses critical gaps in current AI-assisted instructional design regarding transparency, pedagogical foundation, and meaningful human agency through this approach. This research advances the responsible integration of AI in education by providing a concrete, theoretically grounded framework that prioritizes human expertise and educational accountability.","abstract_license":"CC BY 4.0","abstract_license_url":"https://creativecommons.org/licenses/by/4.0/","abstract_source":"https://proceedings.mlr.press/v273/li25a.html","abstract_source_label":"Published version, iRAISE 2025, PMLR 273:94–104","authors":["Hongming Li","Yizirui Fang","Shan Zhang","Seiyon M. Lee","Yiming Wang","Mark Trexler","Anthony F. Botelho"],"bibtex":"https://yiziruifang.com/publication/li-2024-arched/cite.bib","pdf":"https://yiziruifang.com/publication/li-2024-arched/paper.pdf","pdf_source":"https://raw.githubusercontent.com/mlresearch/v273/main/assets/li25a/li25a.pdf","research_topics":["human-centered-ai"],"summary":"Human-centered AI workflow for learning-objective generation, Bloom's taxonomy analysis, and educator-guided instructional design.","title":"ARCHED: A Human-Centered Framework for Transparent, Responsible, and Collaborative AI-Assisted Instructional Design","topics":["Human-centered AI","Instructional design","LLMs in education","Learning objectives","Bloom's taxonomy","Human-AI collaboration"],"url":"https://yiziruifang.com/publication/li-2024-arched/","venue":"Proceedings of the Innovation and Responsibility in AI-Supported Education Workshop (iRAISE 2025), PMLR 273:94–104","year":"2025"},{"abstract":"Learning to defer (L2D) allows prediction tasks to be allocated to a human or machine decision maker, thus getting the best of both’s abilities. Yet this allocation decision depends on a ‘rejector’ function, which could be poorly fit or otherwise misspecified. In this work, we perform uncertainty quantification for the rejector sub-component of the L2D framework. We use conformal prediction to allow the reject to output sets, instead of just the binary outcome of ‘defer’ or not. On tasks ranging from object to hate speech detection, we demonstrate that the uncertainty in the rejector translates to safer decisions via two forms of selective prediction.","abstract_source":"https://openreview.net/pdf?id=TWb9y4PNSW","abstract_source_label":"NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty paper, page 1","authors":["Yizirui Fang","Eric Nalisnick"],"bibtex":"https://yiziruifang.com/publication/fang-2024-learning-workshop/cite.bib","pdf":"https://openreview.net/pdf?id=TWb9y4PNSW","research_topics":["human-ai-decision-making"],"summary":"NeurIPS 2024 workshop paper on conformal uncertainty sets for learning-to-defer rejectors, with abstention and human-model consensus on object and hate-speech detection tasks.","title":"Learning to Defer with an Uncertain Rejector via Conformal Prediction","topics":["Learning to Defer","Conformal Prediction","Uncertainty Quantification","Human-AI Collaboration","Selective Prediction","Abstention"],"url":"https://yiziruifang.com/publication/fang-2024-learning-workshop/","venue":"NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty","year":"2024"},{"abstract":"Inductive conformal predictors (ICPs) are algorithms that are able to generate prediction sets, instead of point predictions, which are valid at a user-defined confidence level, only assuming exchangeability. These algorithms are useful for reliable machine learning and are increasing in popularity. The ICP development process involves dividing development data into three parts: training, calibration and test. With access to limited or expensive development data, it is an open question regarding the most efficient way to divide the data. This study provides several experiments to explore this question and consider the case for allowing overlap of examples between training and calibration sets. Conclusions are drawn that will be of value to academics and practitioners planning to use ICPs.","abstract_license":"CC BY-NC-ND 4.0","abstract_license_url":"https://creativecommons.org/licenses/by-nc-nd/4.0/","abstract_source":"https://arxiv.org/abs/2406.12262v1","abstract_source_label":"arXiv preprint 2406.12262v1, June 18, 2024","authors":["Yizirui Fang","Anthony Bellotti"],"bibtex":"https://yiziruifang.com/publication/fang-2024-investigating/cite.bib","doi":"10.48550/arXiv.2406.12262","pdf":"https://yiziruifang.com/publication/fang-2024-investigating/paper.pdf","pdf_source":"https://arxiv.org/pdf/2406.12262v1","research_topics":["conformal-prediction"],"summary":"How training/calibration splits and overlap affect coverage validity and prediction-set size in inductive conformal classification.","title":"Investigating Data Usage for Inductive Conformal Predictors","topics":["Inductive conformal prediction","Calibration sets","Data splitting","Marginal coverage","Uncertainty quantification","Neural networks"],"url":"https://yiziruifang.com/publication/fang-2024-investigating/","venue":"arXiv preprint arXiv:2406.12262","year":"2024"}],"url":"https://yiziruifang.com/papers.json"}