The paper states that the retrieval corpus contains 14,049 cases aggregated from JuDGE, CAIL2018, CMDL, and LeCaRDv2. However, the public data-preparation pipeline only calls prepare_cail_data.py, and the provided cases_with_feature.json appears to be a small CAIL example. Could you provide the corpus construction script, source case IDs, or the processed 14,049-case corpus used for Table 2?
The paper states that the retrieval corpus contains 14,049 cases aggregated from JuDGE, CAIL2018, CMDL, and LeCaRDv2. However, the public data-preparation pipeline only calls prepare_cail_data.py, and the provided cases_with_feature.json appears to be a small CAIL example. Could you provide the corpus construction script, source case IDs, or the processed 14,049-case corpus used for Table 2?