A curated LLM Security Application related academic papers. All papers are sorted based on the conference name and published year.
💡 Welcome developers or researchers to add more published papers to this list.
Welcome to visit my homepage and Google Scholar.
📬 The cryptocurrency donation address: 0xCC28B05fE858CDbc8692E3272A4451111bDCf700.
| 研究方向 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| 🔐 Security & Crypto | IEEE S&P | ACM CCS | USENIX Security | NDSS | IEEE DSN | SRCS | RAID | CAV |
| 🌐 Networking & Database | SIGMETRICS | ICDE | VLDB | ACM SIGMOD | IEEE INFOCOM | IMC | WWW | |
| 💻 Software Engineering & Programming Language | ICSE | ESEC/FSE | ASE | ACM PLDI | ACM OOPSLA | ISSTA | ACM POPL | SOSP |
| 🧠 Machine Learning | AAAI | ACL | ICML | NeurIPS |
- Recent Advances in Large Language Model Benchmarks Against Data Contamination: From Static to Dynamic Evaluation
- ADVANCES AND CHALLENGES IN FOUNDATION AGENTS FROM BRAIN-INSPIRED INTELLIGENCE TO EVOLUTIONARY, COLLABORATIVE, AND SAFE SYSTEMS
- Foundation Models Defining a New Era in Vision: A Survey and Outlook
- Large Language Models for Blockchain Security: A Systematic Literature Review.
- A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.
- Large language models for software engineering: A systematic literature review.
- Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices.
- Unveiling security, privacy, and ethical concerns of chatgpt.
- Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
- Watermarking Language Models for Many Adaptive Users.
- Prompt Inversion Attack against Collaborative Inference of Large Language Models.
- On the (In)Security of LLM App Stores.
- On the Effectiveness of Prompt Stealing Attacks on In-The-Wild Prompts.
- The Digital Cybersecurity Expert: How Far Have We Come?.
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis.
- PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning.
- Make a Feint to the East While Attacking in the West: Blinding LLM-Based Code Auditors with Flashboom Attacks.
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface.
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks.
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface.
- Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity.
- A Composability Analysis Framework for Web3 Wallet Recovery Mechanisms.
- Analyzing Ad Prevalence, Characteristics, and Compliance in Alexa Skills.
- Prompt Inversion Attack against Collaborative Inference of Large Language Models
- Supporting Human Raters with the Detection of Harmful Content using Large Language Models
- BAIT: Large Language Model Backdoor Scanning by Inverting Attack Target
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- CODEBREAKER: Dynamic Extraction Attacks on Code Language Models
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
- Modifier Unlocked: Jailbreaking Text-to-Image Models Through Prompts
- On Large Language Models’ Resilience to Coercive Interrogation.
- Combing for Credentials: Active Pattern Extraction from Smart Reply.
- DrSec: Flexible Distributed Representations for Efficient Endpoint Security.
- Moderating New Waves of Online Hate with Chain-of-Thought Reasoning in LargeLanguage Models.
- Poisoned ChatGPT Finds Work for Idle Hands: Exploring Developers' Coding Practices with Insecure Suggestions from Poisoned AI Models.
- TROJANPUZZLE: Covertly Poisoning Code-Suggestion Models.
- Transferable Multimoda!Attack on Vision-LanguagePre-Training Models.
- You Only Prompt Once: On the Capabilities of PromptLearning on Large LanguageModels to Tackle ToxicContent.
- SMARTINV: Multimodal Learning for Smart Contract Invariant Inference.
- LLMIF: Augmented Large Language Model for Fuzzing IoT Devices.
- Examining zero-shot vulnerability repair with large language models.
- Analyzing Leakage of Personally Identifiable Information in Language Models.
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions.
- Spinning language models: Risks of propaganda-as-a-service and countermeasures.
- GenderCARE: A comprehensive framework for assessing and reducing gender bias in large language models
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake Detectors
- zkLLM: Zero Knowledge Proofs for Large Language Models
- PLeak: Prompt Leaking Attacks against Large Language Model Applications
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- A Causal Explainable Guardrails for Large Language Models
- ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language Model
- Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack
- PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models
- Legilimens: Practical and Unified Content Moderation for Large Language Model Services
- The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
- PromptFuzz: Prompt Fuzzing for Fuzz Driver Generation.
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models.
- Stealing the Decoding Algorithms of Language Models.
- Large Language Models for Code: Security Hardening and Adversarial Testing.
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
- Protecting intellectual property of large language model-based code generation apis via watermarks.
- Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass.
- Am I Infected? Lessons from Operating a Large-Scale IoT Security Diagnostic Service
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
- LLMmap: Fingerprinting for Large Language Models
- Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
- Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
- APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability Patching
- Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
- StruQ: Defending Against Prompt Injection with Structured Queries
- On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
- Make Agent Defeat Agent: Automatic Detection of Taint-Style Vulnerabilities in LLM-based Agents.
- Rapid Adoption, Hidden Risks: The Dual Impact of Large Language Model Customization.
- PENTESTGPT: An LLM-empowered Automatic Penetration Testing Tool
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.
- EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection.
- Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug Unearthing.
- Lost at c: A user study on the security implications of large language model code assistants.
- CodexLeaks: Privacy Leaks from Code Generation Language Models in GitHub Copilot.
- {Two-in-One}: A Model Hijacking Attack Against Text Generation Models.
- Extracting Training Data from Large Language Models.
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion.
- BumbleBee: Secure Two-party Inference Framework for Large Transformers
- The Philosopher’s Stone: Trojaning Plugins of Large Language Models
- The Skeleton Keys: A Large Scale Analysis of Credential Leakage in Mini-apps
- Safety Misalignment Against Large Language Models
- The (Un)usual Suspects – Studying Reasons for Lacking Updates in WordPress
- CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
- Secure Transformer Inference Made Non-interactive
- Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App Privacy Report
- Provably Unlearnable Data Examples
- YURASCANNER: Leveraging LLMs for Task-driven Web App Scanning.
- LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
- Analysis of the Effect of the Difference between Japanese and English Input on ChatGPT-Generated Secure Codes.
- MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots.
- DeGPT: Optimizing Decompiler Output with LLM.
- DEMASQ: Unmasking the ChatGPT Wordsmith.
- Large Language Model guided Protocol Fuzzing.
- Facilitating Threat Modeling by Leveraging Large Language Models
- Enhancing Static Analysis for Practical Bug Detection: An LLM-Integrated Approach.
- PyDex: Repairing Bugs in Introductory Python Assignments using LLMs.
- Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar.
- Source Code Summarization in the Era of Large Language Models.
- SpecGen: Automated Generation of Formal Program Specifications via Large Language Models.
- SpecRover: Code Intent Extraction via LLMs.
- TOGLL: Correct and Strong Test Oracle Generation with LLMs.
- Template-Guided Program Repair in the Era of Large Language Models.
- Test Intention Guided LLM-based Unit Test Generation.
- The Fact Selection Problem in LLM-Based Program Repair.
- The Seeds of the FUTURE Sprout from History: Fuzzing for Unveiling Vulnerabilities in Prospective Deep-Learning Libraries.
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks.
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models.
- Unleashing the True Potential of Semantic-based Log Parsing with Pre-trained Language Models.
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language Models.
- IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation.
- Hyperion: Unveiling DApp Inconsistencies using LLM and Dataflow-Guided Symbolic Execution.
- Instruct or Interact? Exploring and Eliciting LLMs’ Capability in Code Snippet Adaptation Through Prompt Engineering.
- Intention is All You Need: Refining Your Code from Your Intention.
- InterTrans: Leveraging Transitive Intermediate Translations to Enhance LLM-based Code Translation.
- LLM Assistance for Memory Safety.
- LLM Based Input Space Partitioning Testing for Library APIs.
- LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems.
- LLM-aided Automatic Modelling for Security Protocol Verification.
- LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion.
- LWDIFF: An LLM-Assisted Differential Testing Framework for WebAssembly Runtimes.
- Large Language Models as Configuration Validators.
- Large Language Models for Safe Minimization.
- Leveraging Large Language Models to Detect npm Malicious Packages.
- LiCoEval: Evaluating LLMs on License Compliance in Code Generation.
- LiSSA: Toward Generic Traceability Link Recovery through Retrieval-Augmented Generation.
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language Models.
- Metamorphic-Based Many-Objective Distillation of LLMs for Code-related Tasks.
- Model Editing for LLMs4Code: How Far are We?.
- NIODebugger: A Novel Approach to Repair Non-Idempotent-Outcome Tests with LLM-Based Agent.
- Planning a Large Language Model for Static Detection of Runtime Errors in Code Snippets.
- Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and Defenses.
- ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation.
- RUG: Turbo LLM for Rust Unit Test Generation.
- Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification.
- Reasoning Runtime Behavior of a Program with LLM: How Far Are We?.
- RepairAgent: An Autonomous, LLM-Based Agent for Program Repair.
- Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models.
- RustAssistant: Using LLMs to Fix Compilation Errors in Rust Code.
- SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.
- Search-Based LLMs for Code Optimization.
- A Multi-Agent Approach for REST API Testing with Semantic Graphs and LLM-Driven Inputs.
- Aligning the Objective of LLM-based Program Repair.
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?.
- An Empirical Study on Commit Message Generation using LLMs via In-Context Learning.
- An Exploratory Study of ML Sketches and Visual Code Assistants.
- An LLM-Based Agent-Oriented Approach for Automated Code Design Issue Localization.
- Are LLMs Correctly Integrated into Software Systems?.
- Automated Generation of Accessibility Test Reports from Recorded User Transcripts.
- Automating a Complete Software Test Process Using LLMs: An Automotive Case Study.
- BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks.
- Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference.
- COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge.
- Can an LLM find its way around a Spreadsheet?.
- ChatGPT Inaccuracy Mitigation during Technical Report Understanding: Are We There Yet?.
- ChatGPT-Based Test Generation for Refactoring Engines Enhanced by Feature Analysis on Examples.
- ClozeMaster: Fuzzing Rust Compiler by Harnessing LLMs for Infilling Masked Real Programs.
- Combining Fine-tuning and LLM-based Agents for Intuitive Smart Contract Auditing with Justifications.
- Context Conquers Parameters: Outperforming Proprietary LLM in Commit Message Generation.
- Decoding Secret Memorization in Code LLMs Through Token-Level Characterization.
- Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?.
- DesignRepair: Dual-Stream Design Guideline-Aware Frontend Repair with Large Language Models.
- Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding.
- Feature-Driven End-To-End Test Generation.
- Fixing Large Language Models' Specification Misunderstanding for Better Code Generation.
- HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation.
- Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT
- Fuzz4All: Universal Fuzzing with Large Language Models.
- LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing.
- Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical Study.
- Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT.
- UniLog: Automatic Logging via LLM and In-Context Learning.
- Prompting Is All You Need: Automated Android Bug Replay with Large Language Models.
- Large Language Models for Test-Free Fault Localization.
- Large language models are few-shot testers: Exploring llm-based general bug reproduction.
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context Learning.
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning Libraries.
- GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program Analysis.
- Automated Program Repair in the Era of Large Pre-trained Language Models.
- Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!.
- An Empirical Study of Deep Learning Models for Vulnerability Detection.
- RepresentThemAll: A Universal Learning Representation of Bug Reports.
- Contrabert: Enhancing code pre-trained models via contrastive learning.
- On the robustness of code generation techniques: An empirical study on github copilot.
- Two sides of the same coin: Exploiting the impact of identifiers in neural code comprehension.
- Automated repair of programs from large language models.
- Cctest: Testing and repairing code completion systems.
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language Models.
- Impact of Code Language Models on Automated Program Repair.
- If At First You Don't Suceed, Try, Try, Again ...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software Systems.
- Better Patching Using LLM Prompting, via Self-Consistency.
- Towards Autonomous Testing Agents via Conversational Large Language Models.
- Let's Chat to Find the APIs: Connecting Human, LLM and Knowledge Graph through AI Chain.
- Log Parsing: How Far Can ChatGPT Go?.
- InferFix: End-to-End Program Repair with LLMs.
- Getting pwn'd by ai: Penetration testing with large language models.
- Llm-based code generation method for golang compiler testing.
- Assisting static analysis with large language models: A chatgpt experiment.
- Assess and Summarize: Improve Outage Understanding with Large Language Models.
- Generating realistic vulnerabilities via neural code editing: an empirical study.
- You see what I want you to see: poisoning vulnerabilities in neural code search.
- Not the end of story: An evaluation of chatgpt-driven vulnerability description mappings.
- Understanding Programs by Exploiting (Fuzzing) Test Cases.
- Backdooring Neural Code Search.
- Membership inference attacks against language models via neighbourhood comparison.
- Are you copying my model? protecting the copyright of large language models for eaas via backdoor watermark.
- ReCode: Robustness Evaluation of Code Generation Models.
- Knowledge unlearning for mitigating privacy risks in language models.
1.* Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
1.* Bag of tricks for training data extraction from language models.
1.* Deduplicating training data mitigates privacy risks in language models.
- (Security) Assertions by Large Language Models.
- A Performance-Sensitive Malware Detection System Using Deep Learning on Mobile Devices* A Performance-Sensitive Malware Detection System Using Deep Learning on Mobile Devices.
- PrivacyAsst: Safeguarding User Privacy in Tool-Using Large Language Model Agents.
- CD-VulD: Cross-Domain Vulnerability Discovery Based on Deep Domain Adaptation.
- Software Testing with Large Language Models: Survey, Landscape, and Vision.
- An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation.
- Deep Learning Based Vulnerability Detection: Are We There Yet?.
- On the Value of Oversampling for Deep Learning in Software Defect Prediction.