Designing and Evaluating Capable, Safe, and Trustworthy Generative AI Systems

dc.contributor.advisorGoldstein, Thomas Aen_US
dc.contributor.authorJain, Neel Sindharen_US
dc.contributor.departmentComputer Scienceen_US
dc.contributor.publisherDigital Repository at the University of Marylanden_US
dc.contributor.publisherUniversity of Maryland (College Park, Md.)en_US
dc.date.accessioned2026-07-02T05:42:11Z
dc.date.issued2026en_US
dc.description.abstractGenerative artificial intelligence (GenAI) systems have advanced rapidly throughout the 2020s, reshaping how people interact with technology in everyday life. As these systems are deployed more widely, it has become increasingly important to ensure that they are not only capable but also safe and trustworthy. A guiding perspective of this work is that people should remain at the center of this effort: the methods we develop and the evaluations we design are each motivated by the goal of making generative AI more reliable, more controllable, and more honest for the people who interact with it. Meeting this challenge requires progress across several dimensions, including robustness to adversarial misuse, trustworthy refusal behavior in non-adversarial settings, methods for evaluating models beyond conventional labeled benchmarks, and techniques for improving model capability to ``chat'' under limited supervision. This dissertation investigates these challenges through a study of methods for designing and evaluating generative AI systems. The first part of the dissertation examines the design of safe AI systems in adversarial settings. It studies discrete optimization methods for prompt-based attacks, with a focus on generating effective text inputs that expose model vulnerabilities. In this context, it introduces PEZ, an optimizer for generating adversarial prompts that also supports prompt optimization for broader applications such as image reconstruction and task specification. Building on this formulation, the dissertation presents a series of defenses against gradient-based prompt optimizers and discusses principled approaches for measuring attack success. The dissertation then turns to non-adversarial settings, focusing on the problem of controlling refusal behavior in language models. Because appropriate refusal behavior depends on user preferences, context, and demographic factors such as age, the ability to calibrate refusal behavior is an important aspect of trustworthy AI. Systems are more trustworthy when their boundaries can be adjusted transparently and predictably to suit different users and deployment contexts. To this end, the dissertation introduces meta tokens, referred to as refusal tokens, as a test-time mechanism for controlling specific categories of refusals, and it investigates how to construct data that supports this fine-grained control. The last part of the dissertation focuses on both evaluation and adaptation in settings where conventional supervision is limited. It also examines model capabilities in low-data regimes, where supervised fine-tuning datasets are often small and imperfect, by introducing Noisy Embedding Fine-Tuning (NEFTune), a regularization method that substantially improves response quality and highlights the value of simple interventions for enhancing downstream performance. Finally, the dissertation addresses shortcomings in current benchmarking, as labeled benchmarks are often limited in scale and vulnerable to data contamination. To overcome these issues, it explores evaluation on unlabeled data in a self-supervised setting and introduces a framework for self-sensitivity evaluation, inspired by self-supervised learning, that measures the sensitivity and invariance of language models under transformations of the input text. Collectively, these contributions illustrate that building trustworthy GenAI involves coordinated progress across multiple fronts, by keeping the people who use and are affected by these systems as the context by which to study trustworthy AI.en_US
dc.identifierhttps://doi.org/10.13016/a8ci-y6vl
dc.identifier.urihttp://hdl.handle.net/1903/35867
dc.language.isoenen_US
dc.subject.pqcontrolledComputer scienceen_US
dc.subject.pqcontrolledArtificial intelligenceen_US
dc.subject.pquncontrolledAI Safetyen_US
dc.subject.pquncontrolledLarge Language Modelsen_US
dc.titleDesigning and Evaluating Capable, Safe, and Trustworthy Generative AI Systemsen_US
dc.typeDissertationen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jain_umd_0117E_25980.pdf
Size:
30 MB
Format:
Adobe Portable Document Format