2023
an archive of posts from this year
| Dec 7, 2023 | 系统/工具 · Llama Guard(Meta 内容安全分类器系列) |
|---|---|
| Nov 29, 2023 | 评测集/基准 · MM-SafetyBench(多模态大模型安全评测集) |
| Nov 9, 2023 | 评测集/基准 · SafeBench(FigStep 版,500 条禁忌问题题库) |
| Nov 9, 2023 | 评测集/基准 · FigStep / SafeBench(排版图像越狱评测集) |
| Nov 2, 2023 | 机构 · UK AI Security Institute (AISI) |
| Oct 26, 2023 | 评测集/基准 · ToxicChat(lmsys/toxic-chat) |
| Oct 10, 2023 | 评测集/基准 · MultiJail(Multilingual Jailbreak Challenges 数据集) |
| Oct 10, 2023 | 评测集/基准 · MaliciousInstruct(100 条有害问句 + BERT 打分器) |
| Oct 5, 2023 | 防御机制 · SmoothLLM(随机字符扰动 + 多数投票的越狱防御) |
| Oct 5, 2023 | 评测集/基准 · HEx-PHI(Human-Extended Policy-Oriented Harmful Instruction Benchmark) |
| Aug 18, 2023 | 评测集/基准 · HarmfulQA(declare-lab,Red-Instruct 配套评测集) |
| Jul 27, 2023 | 评测集/基准 · AdvBench(GCG 论文附带的有害行为/有害字符串集) |
| Jul 10, 2023 | 评测集/基准 · BeaverTails(PKU-Alignment 有害性标注语料 / QA-moderation 数据集) |
| Jun 16, 2023 | 防御机制 · Self-Reminder(系统模式自我提醒) |