Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

17 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

CSS-300

🌈 A Multi-Dimensional Benchmark for Decomposing Source-Preference Sycophancy in RAG 🌈

Typing SVG

Made with Love Python RAG License

Stars Forks Issues Last Commit


✨ TL;DR β€” Why CSS-300?

🧠 Does your RAG system trust evidence β€” or does it just trust whichever source it was handed?

CSS-300 is a curated, 300-instance benchmark built to expose and measure source-preference sycophancy: the tendency of a model to defer to a retrieved source's authority instead of independently reasoning about the evidence in front of it.

🎯 Purpose πŸ§ͺ Instances 🧬 Domain πŸ› οΈ Status
Measure sycophancy in RAG 300 curated cases Retrieval-Augmented Generation βœ… Active

🌟 Features at a Glance

πŸ” What's Inside

  • πŸ“¦ dataset.json β€” the full 300-instance benchmark
  • πŸ“Š figures/ β€” every figure from the paper
  • πŸ“ source_data/ β€” raw data behind every figure & table
  • πŸ“œ Fully reproducible pipeline

🎯 What It's For

  • πŸ§ͺ Benchmarking RAG systems
  • πŸ›‘οΈ Testing robustness to biased evidence
  • πŸ”¬ Alignment & interpretability research
  • βš–οΈ Cross-model, cross-pipeline comparison

πŸ—‚οΈ Repository Structure

CSS-300/
β”œβ”€β”€ πŸ“¦ dataset.json          # Complete CSS-300 benchmark dataset
β”œβ”€β”€ πŸ–ΌοΈ  figures/              # Figures included in the manuscript
β”œβ”€β”€ πŸ“Š source_data/          # Source data behind figures & tables
└── πŸ“„ README.md             # You are here πŸ‘‹
File / Directory Description
🧾 dataset.json Complete CSS-300 benchmark dataset
πŸ–ΌοΈ figures/ Figures appearing in the manuscript
πŸ“Š source_data/ Source data to reproduce figures & tables

πŸš€ The Dataset

 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— 
β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β•β•β•     β•šβ•β•β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ•—
β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•‘     β•šβ•β•β•β•β–ˆβ–ˆβ•‘β•šβ•β•β•β•β–ˆβ–ˆβ•‘β•šβ•β•β•β•β•β•šβ•β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘
β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•
 β•šβ•β•β•β•β•β•β•šβ•β•β•β•β•β•β•β•šβ•β•β•β•β•β•β•     β•šβ•β•β•β•β•β•  β•šβ•β•β•β•β•β•  β•šβ•β•β•β•β•β• 

300 carefully curated evaluation instances 🧬 measuring source-preference sycophancy across multiple dimensions.

The benchmark is built for:

  • πŸ§ͺ Benchmarking Retrieval-Augmented Generation systems
  • πŸ›‘οΈ Evaluating model robustness against biased retrieved evidence
  • πŸ“ Measuring source-preference behaviors
  • 🧠 Alignment & interpretability research
  • βš–οΈ Comparative evaluation across language models and retrieval pipelines

πŸ“– For construction methodology, taxonomy, and evaluation protocol, see the accompanying manuscript.


πŸ” Reproducibility

βœ… Included Description
🧾 Complete benchmark dataset
πŸ–ΌοΈ Figures used in the paper
πŸ“Š Source data behind every figure & table

Everything you need to reproduce every analysis and build on top of the benchmark is right here.


πŸ§ͺ Data Generation & Validation

graph LR
    A[πŸ€– AI-Assisted Generation] --> B[πŸ‘₯ Independent Human Annotation]
    B --> C[βœ… Multi-Stage Validation]
    C --> D[🎯 CSS-300 Benchmark]
    style A fill:#F72585,color:#fff,stroke:#333
    style B fill:#7209B7,color:#fff,stroke:#333
    style C fill:#3A0CA3,color:#fff,stroke:#333
    style D fill:#00C9A7,color:#fff,stroke:#333
Loading

The benchmark was generated via an AI-assisted pipeline, then independently annotated and validated by the research team, with multiple stages of quality-assurance review.


πŸ”’ Ethics & Privacy

No PII No Personal Data Research Only

This repository does not contain:

  • ❌ Personal data
  • ❌ Sensitive information
  • ❌ Personally identifiable information (PII)
  • ❌ Human subject data

Intended exclusively for scientific research on RAG systems and model evaluation. πŸ”¬


πŸ“š Citation

If CSS-300 powers your research, please cite:

@article{css300,
  title={CSS-300: A Multi-Dimensional Benchmark for Decomposing Source-Preference Sycophancy in Retrieval-Augmented Generation},
  author={Authors},
  year={2026}
}

✏️ Update with final publication details once available.


πŸ“„ License

Please refer to the repository's LICENSE file for licensing information.


πŸ‘₯ Contributors


πŸ“¬ Contact

Questions, issues, or collaboration ideas? πŸŽ‰

  • πŸ› Open an issue in this repository
  • βœ‰οΈ Reach out to the corresponding authors listed in the manuscript

⭐ If CSS-300 helps your research, consider starring the repo! ⭐

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors