					Trusted AI Challenge


Team: EthosIntel, University of Wisconsin Madison [Benjamin Afflerbach, Ross Klein, Arundhati Singh, Aditya Rawat, Edison Chiu]


How would your team measure success in the Competition as a model developer team?

Our team is driven by the motivation to welcome a new generation of safer AI models. This challenge has given us a valuable opportunity to explore the safety and security of AI models by studying various attacks and defenses through numerous profound research papers. We believe that Artificial Intelligence should foster growth, and it is crucial to build models that protect against the exploitation of personal or sensitive information.
We rely heavily on an agentic approach that iteratively processes prompts. Our procedure involves sending the prompt through a code-generation model and a maliciousness critic, who then sends their decisions to a fine-tuned judge model. The judge model validates the prompt as safe or unsafe, and if deemed safe, it is returned to the code-generation model to generate secure and efficient code. If the generated code doesn’t meet our safety and efficiency standards, it is fed back into the code-generation model for further refinement.
Our goal is to enhance the scalability of this model, ensuring both speed and ease of use. The key to building a safe AI model is to develop an efficient security system that works across different models and serves as a foundational evaluation framework. Our fine-tuned agents will focus on generating secure yet efficient code that can be scaled across a wide range of platforms. This will allow companies to seamlessly integrate our model into their systems, ensuring both safety and usability.

Describe, in detail, your scientific approach along with a related system architecture.

Section 1: Multi-Agent System 
To enhance the effectiveness of AI moderation, we propose a sophisticated Multi-Agent System that integrates fine-tuning of Large Language Models (LLMs) with a multi-layered critique framework. This system is designed to address the nuanced challenges of code evaluation and security assessment through a structured, iterative process.

System Overview:

Code Generation Model: The process begins with a Code Generation Model, which receives the initial prompt and generates code based on its interpretation. This model serves as the starting point for producing code outputs. The code generation model is expected to keep a chat history. Prompts identified as malicious can be removed from the chat history.
Maliciousness Critic: Concurrently, the prompt is also passed to the Maliciousness Critic. This agent, aligned using prompt engineering assesses the prompt and flags potential risks or vulnerabilities with a high degree of sensitivity. Our goal is to have the maliciousness critic provide a security oriented critique for the Judge, regardless of how benign the input is. This is to provide accountability to the code generation model, regardless of if it was compromised or not. The Critic will not keep a chat history.
Judge LLM: Both the Code Generation Model’s output and the Maliciousness Critic’s assessment are passed to the Judge LLM. The judge will make an informed decision of whether the code is malicious, leveraging its contextual understanding and implicit decision-making capabilities to determine whether the code is secure or poses any risks. The Judge will not keep a chat history.
Agent Decision: If the Judge LLM identifies potential issues, the prompt is rejected, and the rejection output will be returned, and the loop stops. Else, the loop waits for the feedback loop occurring through the code generation model, safety critic, and the helpful critic to complete before safe and efficient code is returned to the user. 
Critique Phase – Iterative Code Generation: The code generated by the code generation model is concurrently passed through two more agents. Once we have determined that the code requested is not malicious, we must determine if the generated code poses any security risks to the user. The goal of this phase is to ensure the security and accuracy (to the user question) of the code.-
Safety Critic: This agent evaluates the safety of the generated code, checking for potential security threats and compliance with safety standards. The agent will provide a short description critiquing the code.
Helpful Critic: Simultaneously, this critic assesses the quality of the code, ensuring that it meets the desired standards of functionality and effectiveness. The agent will provide a short description critiquing the code.
 Iterative Feedback Loop: The code is then cycled through the Code Generation Model several times, incorporating feedback from the Safety and Helpful Critics to enhance both its security and quality. During the lifetime of a code snippet, the critics are expected to maintain a chat history of their agentic interactions. The history of the critics can be reset afterwards.
User Delivery: After passing through these rigorous evaluations and iterative refinements, the final code is made available to the user if the judge LLM accepts the prompt, ensuring that it is both secure and high-quality.
Advantages:

This Multi-Agent System effectively addresses key challenges in AI moderation by employing fine-tuning and specialized agents tailored for specific tasks. The fine-tuning of the LLMs enhances their ability to identify and critique both benign and potentially malicious code efficiently. The multi-layered critique framework reduces the area of attack of an adversary on the model and effectively reduces all attacks to template attacks. Once the code generation model expands any potential malicious prompting into a code snippet, the nature of the task should be readily apparent to the Maliciousness Critic agent and the Judge model.


Section 2: Fine-Tuning the Judge

In order to improve the Judge model’s proficiency in comprehending review tasks and adhering to task instructions, we will fine-tune the model. We do not expect the fine-tuning to improve the models inherent ability to distinguish between malicious and non-malicious questions, but we do expect the fine-tuning to make the model more reliable and perform closer to its capabilities. 

System Overview:

Input Datasets: The pipeline starts with two types of data:
Code Generation Datasets: There are a variety of code generation datasets available. Ideally, we would like to choose the same, or similar datasets to those that the code generation model itself was trained on.
Jailbreak Templates: We have compiled a variety of jailbreak strategies to randomly select from. These strategies all come from published papers.
Modified AART pipeline: Taking inspiration from the strategy used by Radharapu, Bhaktipriya, et al. we will devise a pipeline to randomly modify instructions to both be malicious and employ an adversarial attack strategy.
The pipeline splits the code generation dataset in a 1:1 ratio.
For the portion processed by AART, it retrieves questions from the dataset and applies a random jailbreak template to generate unique malicious questions
Non-Malicious Questions: The other half of the dataset (non-malicious questions) is kept as is.
Code Generation: Using the code generation model, we will generate outputs for each of the modified malicious instructions.
Critique Generation: The model’s code generation outputs (malicious and non-malicious) are critiqued using a “maliciousness critic” to assess the potential harmfulness of the generated content
Modified Dataset for Fine-Tuning: The outputs of the pipeline are assembled into a new dataset which can be formatted with a prompt template for fine-tuning. The final dataset will be a mix of both malicious and non-malicious queries. The dataset will have the following format
Instructions: The dataset will include the input prompt that generated the following code.
Generated code: The dataset will  include the output of the above instructions
Maliciousness Label: The dataset will annotate each query, identifying whether it was modified to be malicious or not.

Advantages:

The Fine-Tuning pipeline effectively trains the model on its expected input in a robust manner. This step will improve the models performance on unseen tasks without the need for few-shot prompting. We plan on using low-rank adaptation (LoRa), a method used for parameter-efficient fine-tuning (PEFT) of large language models, to finetune the Judge LLM. This approach is more efficient computationally and allows for a smaller dataset while producing similar if not better results over full fine-tuning methods.


Got it — focusing just on the **Model Pipeline diagram** (page 4), here's a clean summary you can pair with the full text:

---

**Model Pipeline Diagram Summary (can't paste the image so we summarize it instead for the md)**

The diagram illustrates two distinct pipelines: the **inference-time Model Pipeline** and the **training-time Fine-Tune Pipeline**.

**Model Pipeline (top):**
A user question enters two parallel paths simultaneously — it goes to the **Code Generation Model** and the **Maliciousness Critic** at the same time. The Maliciousness Critic is a deliberately strict base model that explains how any generated code *could* be used maliciously, regardless of whether it actually is. Both outputs feed into the **Judge LLM** — a fine-tuned model — which makes the accept/reject decision. If rejected, a rejection output is returned immediately. If accepted, the generated code is passed through two more agents running in parallel — a **Safety Critic** and a **Helpful Critic** — which iteratively refine the code over two cycles before the final output is returned to the user.

**Fine-Tune Pipeline (bottom):**
The training pipeline starts with a code generation dataset, split 50/50. Half is kept as clean, non-malicious data. The other half is processed through **AART**, which applies randomly selected jailbreak templates (drawn from published papers including DeepInception, FuzzLLM, Jailbroken, and roughly 78 others) to generate adversarial variants. Both halves are passed through the code generation model and critiqued by the Maliciousness Critic. The resulting dataset — containing generated code, critiques, and maliciousness labels — is used to fine-tune the Judge LLM.

---
What is novel about your approach?

The novelty of our approach primarily lies in the strategic integration of a multi-agent system that not only tackles the security concerns of AI-generated code but also emphasizes its overall functionality and quality. 

Traditional moderation systems typically rely on a single model or agent to assess input for potential vulnerabilities, which often results in limited coverage of the complexities involved in both security and functionality. Our approach, however, innovates by introducing a suite of specialized agents—each tailored to perform distinct tasks through the use of advanced prompt engineering.

A key element that elevates our system is the integration of the Judge LLM and surrounding architecture. Instead of preventing the code generation model from producing malicious code, we embrace its functionality and instead moderate downstream. We implement a Maliciousness Critic, meant to provide robust feedback, even being too strict, on the generated code to ensure that no malicious features are hidden in the generated response. The maliciousness critic will provide the necessary context to the Judge which will make an informed decision on whether the generated code is malicious or not. To avoid providing false positives, the Judge will be carefully aligned so that it will weigh the critiques against its own reasoning, ensuring that the final decision is balanced and aligned with security protocols. 

The Judge will also undergo rigorous fine-tuning so that it will perform its goal to the best of the models capabilities. Our research into fine-tuning has led us to believe that fine-tuning is exceptionally useful when it comes to aligning the behaviors of LLMs with a specific task. We expect fine-tuning to greatly enhance the reliability and overall effectiveness of the Judge model. Specifically, fine-tuning will limit the situations in which the model will misunderstand or have difficulty with a task, increasing the opportunity for proper assessment. Although, we found contrary evidence regarding whether fine-tuning can enhance the upper limit of assessment capabilities of the model.

The design of the Judge LLM architecture also reduces the surface area of attack drastically. The goal of an adversary is to hide malicious instructions from the models safety measures. 
By allowing the code generation model to produce malicious code, we unravel the hidden malicious intent in the adversaries attack, allowing much easier identification. We expect the code generation model to become compromised, as it is the only vector of attack for an adversary. The only way a downstream agent of the code generation model can be attacked is if an adversary were to fool the code generation model into producing an adversarial prompt itself. Furthermore, we expect any generative/fuzzing approaches to be rendered ineffective as there is no continuous history for the Malicious Critic agent or the Judge. Training gaps attacks are still a possibility, such as low-resource language attacks; it is unclear how the architecture will handle these cases.

What truly sets our system apart is the balance it achieves between security and functionality. In traditional systems, there is often a trade-off between ensuring security and maintaining the operational efficiency of the code. However, our system does not prioritize one over the other. Instead, it introduces a modular and layered evaluation process that rigorously assesses both aspects. By involving the INDICT Safety Critic and Helpful Critic framework to enhance the functionality of the code, we ensure that the code is not only safe and secure but also effective in its intended function. Furthermore, the adaptability of the system allows for easy scalability, enabling the integration or removal of agents or critics as required, making it a highly versatile solution.

In summary, the combination of fine-tuning, Judge architecture, and a multi-layered critique framework positions our approach as a novel and sophisticated solution for generating AI-generated code that meets high standards of security, functionality, and reliability. By bridging the gap between static evaluations and dynamic, multi-agent assessments, our system offers a more holistic, adaptable, and robust approach to code evaluation and generation in AI applications.

How do you think your work will impact the field of Responsible AI and code generation?

Our team is dedicated to making meaningful contributions to responsible AI and code generation by concentrating on safety, transparency, and effectiveness. We hope to see our efforts contribute to the monumental advancement of AI. We feel as if we are on the horizon of a great change in society and it would be an honor to contribute to it. Our team is setting high standards for safety, transparency, and effectiveness in AI-generated code. Our approach aims to create secure, reliable, and accountable AI technologies that have a positive impact on the field.



Please provide a summary of technical work and research (relevant to your proposed architecture), yours or others’, that you will leverage and how.

Fine tuning and alignment:

Le, Hung, et al. "INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness." arXiv preprint arXiv:2407.02518 (2024).
Zeng, Wenjun, et al. "ShieldGemma: Generative AI Content Moderation Based on Gemma." arXiv preprint arXiv:2407.21772 (2024).
Balne, Charith Chandra Sai, et al. "Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications." arXiv preprint arXiv:2404.13506 (2024).
Huang, Hui, et al. "An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge models are task-specific classifiers." arXiv preprint arXiv:2403.02839 (2024).
Lu, Junyi, et al. "LLaMA-Reviewer: Advancing code review automation with large language models through parameter-efficient fine-tuning." 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 2023.
Wei, Jason, et al. "Finetuned language models are zero-shot learners." arXiv preprint arXiv:2109.01652 (2021).
Ouyang, Long, et al. "Training language models to follow instructions with human feedback." Advances in neural information processing systems 35 (2022): 27730-27744.
Schulhoff, Sander, et al. "The Prompt Report: A Systematic Survey of Prompting Techniques." arXiv preprint arXiv:2406.06608 (2024).

Justification and Pitfalls with Fine Tuning in Judge Models
Fine tuning with moderation has been shown to be ineffective (4) but that doesn’t mean it doesn’t have its place in a moderation model. We know from foundational papers in the fine-tuning field (6) (7) that fine tuning is effective in improving performance on unseen NLP tasks. “Though finetuning is very effective on tasks naturally verbalized as instructions and is less effective on tasks directly formulated as language modeling, where instructions would largely be redundant”  (6). Given the evidence that (4) presents on how finetuning has been shown to be ineffective on moderation tasks, moderation appears to be a language modeling task and not an instructable task. Paper (5) supports this conclusion despite using fine-tuning as a strategy in developing a code generation model. Paper (5) employs fine-tuning to enhance “the models proficiency in comprehending code review tasks and adhering to task instructions” and to aid “the model in better interpreting user intentions and following instructions”. This is an appropriate use of fine-tuning. Though, it is clear that beyond instruction and goal alignment, fine-tuning cannot increase the ability of an LLM to moderate. PEFT is a more efficient and effective way to finetune a model (3). Paper (5) uses PEFT, specifically the LoRa training strategy to fine-tune. LoRa (low-rank adaptation) freezes the original model and fine tunes a separate set of weights which are added to the original parameters. LoRa reduces the number of parameters by transforming the model into a lower-rank dimension.


Prompting Strategies and Best Practices
There is a plethora of research on prompt engineering approaches being used and researched across thousands of papers (8). These approaches can be broken down into 50 or so categories. Previous work in prompt engineering with agents on code generation models shows the use of self-criticism as a viable strategy (1). Self-criticism and chain of thought seem to be the optimal way of approaching nuanced judgment tasks. Paper (1) specifically uses the self-refine criticism method to iteratively improve the original output based on agentic feedback. “Self-Refine has demonstrated improvement across a range of reasoning, coding, and generation tasks” (8). Answer engineering is the process of extracting precise answers from LLM outputs (8). Fine-tuning is extremely effective at embedding good answer formats into the LLMs before any answer engineering is done (5). The final mile for a fine-tuned LLM still comes from prompting. You must instruct the LLM to output text in a specific format so it can be extracted using expression capturing like REGEX.




Jailbreaking Overview

Xu, Zihao, et al. "LLM Jailbreak Attack versus Defense Techniques--A Comprehensive Study." arXiv preprint arXiv:2402.13457 (2024).
Abdali, Sara, et al. "Can LLMs be Fooled? Investigating Vulnerabilities in LLMs." arXiv preprint arXiv:2407.20529 (2024).
Yang, Zeyu, et al. "Assessing Adversarial Robustness of Large Language Models: An Empirical Study." arXiv preprint arXiv:2405.02764 (2024).


Techniques and Considerations for Jailbreak attack

Jailbreak attacks on large language models (LLMs) demonstrate varying success rates depending on the model's architecture and defense mechanisms. Template-Based Attacks are often the easiest to execute and highly effective against less advanced models, while Generative Attacks, which use complex algorithms, tend to bypass more sophisticated defenses (1). Paraphrasing and spoofing attacks have the potential to be severe vulnerabilities (2). Paraphrasing attacks are when an adversary modifies the input text to the LLM using a paraphrase model, potentially evading safeguards while holding the same semantic meaning (2). Spoofing attacks target the LLMs chat history by imitating the LLM and manipulating the chat history with compromising LLM outputs, resulting in jailbroken generation (2). Training Gap Exploits target specific vulnerabilities in a model's safety training. Overall, attack success is influenced by the model's robustness and the complexity of the attack strategy employed. White box attacks are generally less effective than black-box attacks (1). White-box attacks generally involve access to internal metrics such as loss metrics, making the attack impractical in many cases. Some models tend to perform better against template-based attacks because of the model’s training on diverse datasets that include common adversarial attacks (1). However, on the contrary, generative forms of attacks are usually more successful than other forms of attacks against models that lack robust generative defenses, indicating that the complexity of these generative forms of attacks can overwhelm the models’ traditional defenses or safeguards (1). 





Techniques and Considerations for Jailbreak Defense

Jailbreaking defense techniques differ in their effectiveness depending on the attack type and the model's approach to security. Self-Processing defenses work well against simpler attacks, while Input Permutation defenses enhance model robustness by varying inputs (1). There are three main categories of jailbreaking mitigation strategies: self-processing defense, additional helper defense, and Input permutation defense (2). “Self-Processing Defenses, which rely exclusively on the LLM’s own capabilities; Additional Helper Defenses, which require the support of additional algorithms or auxiliary LLMs for verification purposes; and Input Permutation Defenses, which manipulate the input prompt and verify with the target LLMs mutliple times to detect and counteract malicious requests aimed at exploiting gradient-based vulnerabilities.” (2) Self-processing defenses includes techniques like embedding the user query within a prompt and further agentic approaches. Additional helper approaches include employing auxiliary LLMs to check or modify the query, performing calculations on or manipulating token-level perplexity. Input permutation involves partial deletion of input content, modifying prompts by swapping or adding, or random dropping of input. Models that rely on self-processing defenses often show a higher rate of success against simpler forms of attacks, but fail to defend themselves against more sophisticated forms of attacks like training gap exploits (1). On the contrary, models employing external helper defenses can better handle complex attacks but may potentially suffer from higher false positives (may classify more benign prompts as harmful) (1).

Data vulnerabilities
Training data is a critical vulnerability point in LLMs in pretraining and fine-tuning. “Given the remarkable effectiveness of data poisoning, there arises a necessity to mitigate the potential harm it can cause.” (2). This goes for accidental data poisoning as well. Data with bad distributions of information, such as not including enough adversarial attacks, can lead to exploitable vulnerabilities. Mitigation strategies include validating training data from trusted sources, sanitizing and preprocessing data to remove toxic examples.

Jailbreaking techniques
Radharapu, Bhaktipriya, et al. "Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications." arXiv preprint arXiv:2311.08592 (2023).
Kang, Daniel, et al. "Exploiting programmatic behavior of llms: Dual-use through standard security attacks." 2024 IEEE Security and Privacy Workshops (SPW). IEEE, 2024.
Yao, Dongyu, et al. "Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models." ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024.
Wei, Alexander, Nika Haghtalab, and Jacob Steinhardt. "Jailbroken: How does llm safety training fail?." Advances in Neural Information Processing Systems 36 (2024).
Du, Yanrui, et al. "Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak." arXiv preprint arXiv:2312.04127 (2023).
Liu, Yi, et al. "Jailbreaking chatgpt via prompt engineering: An empirical study." arXiv preprint arXiv:2305.13860 (2023).

