CS 684: Trustworthy and Responsible AI

Instructor: Eugene Bagdasarian
TA: Ali Naseh, Dzung Pham
Time: Mon & Wed 10:00AM - 11:15AM
Location: Computer Science Building (CSW) 150
Office hours: Eugene: TBD, CSL 451 | Ali: By appointment, CSL 347 | Dzung: By appointment, CSL 347

In the era of intelligent assistants, autonomous agents, and self-driving cars, we expect AI systems to not cause harm and withstand adversarial attacks. In this course, you will learn advanced methods of building AI models and systems that mitigate privacy, security, societal, and environmental risks. We will go deep into attack vectors and what type of guarantees current research can and cannot provide for modern generative models. The course will feature extensive hands-on experience with model training and regular discussion of key research papers. Students are required to have taken NLP, general ML, and security classes before taking this course.

Expectations

  • Required reading, attendance, and participation
  • Each group: weekly presentation + code for assignments
  • Group research project

Grading Breakdown

Component Percentage Details
Attendance 10% Allowed to miss any 4 classes
Assignments (slides + report + code) 40% 2 total (20% each), allowed 2 late day per assignment.
Final Project 50%
  • 1-page Proposal (10%)
  • Mid-Semester Presentation (5%)
  • Final Report (20%)
  • Final Presentation (15%)
(Optional) bonus up to 5% Active participation, excellent code implementation, slide efforts

Syllabus: Weekly Schedule

Fall 2026 Class Schedule
Week Class # Date Topic Notes Links/Slides/Assignments
Week 1 1 Wed, Sep 9 Introduction First Day of Classes Assignment 0: Startup ideas
No reading
Week 2 2 Mon, Sep 14 Overview of Privacy and Security Slides
📖 Paper 1: Towards the Science of Security and Privacy in Machine Learning
3 Wed, Sep 16 Privacy. Membership Inference Attacks Slides
📖 Paper 1: Membership Inference Attacks From First Principles
📖 Paper 2: Membership Inference Attacks against Machine Learning Models
Week 3 4 Mon, Sep 21 Privacy. Training Data Attacks Last day to add/drop (Grad) Assignment 1 Release (Due 10/2)
📖 Paper 1: Extracting Training Data from Large Language Models
📖 Paper 2: Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
📖 Paper 3: Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
📖 Paper 4: Imitation Attacks and Defenses for Black-box Machine Translation Systems
5 Wed, Sep 23 Privacy. Federated Learning
📖 Paper 1: Communication-Efficient Learning of Deep Networks from Decentralized Data
📖 Paper 2: Advances and Open Problems in Federated Learning
Week 4 6 Mon, Sep 28 Privacy. Differential Privacy, Part 1. Basics Assignment 2 Release (Due 10/9)
📖 Paper 1: Deep Learning with Differential Privacy
📖 Paper 2: Scaling Laws for Differentially Private Language Models
7 Wed, Sep 30 Privacy. Differential Privacy, Part 2. In-Context Learning, Private Evolution
📖 Paper 1: Differentially Private Synthetic Data via Foundation Model APIs 1: Images
📖 Paper 2: Learning Differentially Private Recurrent Language Models
Fri, Oct 2 Assignment 1 due
Week 5 8 Mon, Oct 5 Privacy. Data Analytics, PII Filtering with LLMs Assignment 3 Release (Due 10/16)
📖 Paper 1: Beyond Memorization: Violating Privacy Via Inference with Large Language Models
📖 Paper 2: Can Large Language Models Really Recognize Your Name?
9 Wed, Oct 7 Privacy. Contextual Integrity
📖 Paper 1: Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
📖 Paper 2: AirGapAgent: Protecting Privacy-Conscious Conversational Agents
Fri, Oct 9 Assignment 2 due
Week 6 10 Mon, Oct 12 No class (Indigenous People's Day) Assignment 4 Release (Due 10/23)
No paper reading
11 Wed, Oct 14 Privacy. Guest Lecture: TBD
No paper reading
Fri, Oct 16 Assignment 3 due
Week 7 12 Mon, Oct 19 Mid-Semester Project Presentations Instructions
13 Wed, Oct 21 Mid-Semester Project Presentations (cont.)
Fri, Oct 23 Assignment 4 due
Project Proposal due
Week 8 14 Mon, Oct 26 Security. Jailbreaks + Prompt injections Assignment 5 Release (Due 11/06)
📖 Paper 1: Universal and Transferable Adversarial Attacks on Aligned Language Models
📖 Paper 2: Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
15 Wed, Oct 28 Security. Guest Lecture (Adversarial ML)
Fri, Oct 30
Week 9 16 Mon, Nov 2 Security. Adversarial Examples in Multi-modal systems Assignment 6 Release (Due 11/13)
📖 Paper 1: Are aligned neural networks adversarially aligned?
📖 Paper 2: Self-interpreting Adversarial Images
17 Wed, Nov 4 Security. Poisoning and Backdoors Last day to drop (Grad)
📖 Paper 1: How To Backdoor Federated Learning
Fri, Nov 6 Assignment 5 due
Week 10 18 Mon, Nov 9 Security. Watermarks Assignment 7 Release (Due 11/20)
📖 Paper 1: SoK: Watermarking for AI-Generated Content
19 Wed, Nov 11 No class (Veterans' Day)
Fri, Nov 13 Assignment 6 due
Week 11 20 Mon, Nov 16 Security. Alignment Attacks Assignment 8 Release (Due 11/29)
📖 Paper 1: Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Week 12 22 Wed, Nov 18 Security. Principled Defenses
📖 Paper 1: Defeating Prompt Injections by Design
📖 Paper 2: Contextual Agent Security
Fri, Nov 20 Assignment 7 due
21 Mon, Nov 23 Environmental. Overthinking
📖 Paper 1: OverThink: Slowdown Attacks on Reasoning LLMs
23 Tue, Nov 24 Societal. Propaganda, Misinformation, and Deception Thanksgiving recess begins after last class
📖 Paper 1: Propaganda-as-a-service
Sun, Nov 29 Assignment 8 due
Week 13 24 Mon, Nov 30 Societal. Fairness, Ethics.
📖 Paper 1: Differential Privacy Has Disparate Impact on Model Accuracy
25 Wed, Dec 2 Guest Lecture: TBD Final Report/Poster Instructions
No paper reading
Week 14 26 Mon, Dec 7 Guest Lecture: TBD
No paper reading
27 Wed, Dec 9 Guest Lecture: TBD
No paper reading
Fri, Dec 11 Poster due
Week 15 28 Mon, Dec 14 Final Project Poster Session
Week 16 28 Wed, Dec 23 Final Project Report Due (Last day of semester)

Assignments

All Assignments Overview

  1. Build synthetic data + Show attacks (MIA or data extraction) → reconstruction
  2. Implement Federated Learning
  3. Implement Differential Privacy and Private Evolution
  4. Contextual Integrity & PII Filtering: PII filtering/extraction; contextual integrity → airgap, context hijacking
  5. Jailbreaking and prompt injections
  6. Multi-modal attacks
  7. Backdoors and watermarking
  8. Alignment attacks + RLHF

Assignment Process

  1. Create your repository: Each team must create a repository on GitHub.
  2. Share access: Add the teaching staff as collaborators.
  3. Roles:
    • One lead author writes the report and conducts the main experiments.
    • Other team members advise and consult.
    • The lead author receives 80% of the grade, other members receive 20%.

Structure of Each Assignment

  • Reading Report: Summarize, critique, and connect the assigned papers to your project.
    Include key discussion points from class as well.
  • Code: Implement the required attacks/defenses, include documentation and results.
  • Presentation: Prepare a short slide deck summarizing your findings.

Deadlines & Presentations

  • Due: Slide deck, reading report, and code are due Friday of that week, 11:59 PM EST.
  • Presentation: Happens at the beginning of the following week’s class.

Course Policies

Course Grade Scale

Grade Range
A 93-100
A- 90-92
B+ 87-89
B 83-86
B- 80-82
C+ 77-79
C 73-76
F 0-72

Notes on AI Use

You may use AI tools to help with reading or drafting, but you must fully understand the material and be able to explain it clearly to your teammates. The goal is to enrich group learning and class discussion—not just to generate text. You need to provide the “How I used AI” section in your report.

Late Policy

Each assignment includes two late days (24 hours) that may be used without penalty. Late days do not accumulate across assignments—unused late days expire. Assignments submitted beyond the allowed late day will not be accepted unless prior arrangements are made due to documented emergencies.

Nondiscrimination Policy

This course is committed to fostering an inclusive and respectful learning environment. All students are welcome, regardless of age, background, citizenship, disability, education, ethnicity, family status, gender, gender identity or expression, national origin, language, military experience, political views, race, religion, sexual orientation, socioeconomic status, or work experience. Our collective learning benefits from the diversity of perspectives and experiences that students bring. Any language or behavior that demeans, excludes, or discriminates against members of any group is inconsistent with the mission of this course and will not be tolerated.

Students are encouraged to discuss this policy with the instructor or TAs, and anyone with concerns should feel comfortable reaching out.

Academic Integrity

All work in this course is designated as group work, with shared responsibility among members. While assignments will be submitted jointly and receive a group grade, each member is expected to contribute meaningfully and to track individual contributions within the group.

Collaboration within your group is encouraged and expected. You may discuss ideas, approaches, and strategies with others, but all written material, whether natural language or code, must be the original work of your group. Copying text or code from external sources without proper attribution is a violation of academic integrity.

This course follows the UMass Academic Honesty Policy and Procedures. Acts of academic dishonesty, including plagiarism, unauthorized use of external work, or misrepresentation of contributions, will not be tolerated and may result in serious sanctions.

If you are ever uncertain about what constitutes appropriate collaboration or attribution, please ask the instructor or TAs before submitting your work.

Accommodation Statement

The University of Massachusetts Amherst is committed to providing an equal educational opportunity for all students. If you have a documented physical, psychological, or learning disability on file with Disability Services (DS), you may be eligible for reasonable academic accommodations to help you succeed in this course. If you have a documented disability that requires an accommodation, please notify me within the first two weeks of the semester so that we may make appropriate arrangements. For further information, please visit Disability Services.

Title IX Statement (Non-Mandated Reporter Version)

In accordance with Title IX of the Education Amendments of 1972 that prohibits gender-based discrimination in educational settings that receive federal funds, the University of Massachusetts Amherst is committed to providing a safe learning environment for all students, free from all forms of discrimination, including sexual assault, sexual harassment, domestic violence, dating violence, stalking, and retaliation. This includes interactions in person or online through digital platforms and social media. Title IX also protects against discrimination on the basis of pregnancy, childbirth, false pregnancy, miscarriage, abortion, or related conditions, including recovery.

There are resources here on campus to support you. A summary of the available Title IX resources (confidential and non-confidential) can be found at the following link: https://www.umass.edu/titleix/resources. You do not need to make a formal report to access them. If you need immediate support, you are not alone. Free and confidential support is available 24/7/365 at the SASA Hotline 413-545-0800.