Early-exit neural networks (EENNs) let a model stop computing as soon as it is confident enough in a prediction, cutting inference cost without running every layer. What decides when to stop is the routing strategy. The field has produced many of them, but they are almost always evaluated in isolation, on different models, training setups, and datasets, which makes it hard to tell whether a strategy is genuinely better or simply paired with a better model.
I am building a unified framework for comparing routing strategies across approaches. Its core is a fittability protocol built around one question: given a routing strategy R and a model M, can R be applied to M, given the signals, architecture, and other requirements R needs? Strategies that fit the same model can be compared in a controlled setting where the routing strategy is the only independent variable, and are benchmarked against task-specific oracles.
For strategies that cannot be fit to a shared model, I am defining a fallback protocol that compares model–routing pairs on shared datasets while accounting for differences in model and training, so routing performance can still be separated from model training. Because what counts as a good exit differs across image classification, object detection, and text generation, I am also formulating upper-bound oracles tailored to each task's early-exit priorities.
Most routing strategies rely on confidence-like signals, so I am also formalizing a notion of trustworthiness for early-exit networks, using the calibration of those exit signals as the defining criterion.
Ongoing — framework and protocols are under active development.
In this project, funded by two cycles of the NSF Computing Alliance of Hispanic-Serving Institutions REU program, I investigated autonomous computer-use agents (ACUAs) — systems powered by large language models that can operate a computer end-to-end. Unlike traditional chatbots, ACUAs navigate interfaces, execute tasks, and make independent decisions, raising important questions about their reliability and security.
I designed and introduced one of the first systematic evaluation frameworks for ACUAs, testing agents from OpenAI, Anthropic, and open-source projects across five task domains of increasing complexity, adapting principles of an HCI IBM UI/UX quantitative assessment to measure complexity. The study identified two classes of agents: full computer access and browser-based agents.
Performance was measured with a seven-factor rubric assessing accuracy, adaptability, efficiency, robustness, security, relevance, and consistency. Quantitative data such as completion rates, time, failed interactions, and remediation percentages were collected for evaluation.
Findings revealed significant limitations: full computer-access agents often failed due to hallucinations, navigation errors, and unauthorized system changes, while browser-based agents achieved higher success rates but still showed vulnerabilities to prompt injection and inconsistent security awareness.
These results currently guide the development of an ACUA, integrating multi-agent orchestration, machine learning, RAG, and security frameworks including access control and prompt verification. Potential approaches for addressing LLM limitations, including chain-of-thought (CoT) reasoning, are under investigation to improve decision making.
This work has been recognized with multiple honors, including a GMiS Student Poster Scholarship, an NSF LSAMP Scholarship, a second NSF CAHSI REU award, and a scholarship with second place at the international WiCyS Student Poster Competition.
Funded by the NSF Smart Cities REU program, I investigated large language model optimization techniques for analyzing unstructured construction safety data to enhance urban infrastructure development. LLM-pipeline optimization enables automated extraction of critical insights from accident reports, raising important questions about scalability and real-world implementation in smart city frameworks.
Project I: As the computer scientist in a team of civil engineers, I developed novel LLM-pipeline optimization methods to analyze construction accident reports, creating a prototype that processes unstructured safety data and extracts meaningful insights. My work included researching optimized RAG and machine learning pipelines, and evaluating strategies to improve efficiency while enhancing narrative tone extraction.
Project II: I collaborated on NDA research focused on machine learning for system analysis, working on techniques designed to enhance LLM performance and improve the accuracy of pattern extraction as the only computer scientist on the team.
I have been invited to continue collaborating as a research assistant. Project I has evolved quickly and I now lead the LLM-driven analytical methods, contributing to manuscript development and gaining experience integrating quickly into a new interdisciplinary team.