Academic

Academic

Academic · 약 1분

MAC: Multi-Agent Constitution Learning

arXiv:2603.15968v1 Announce Type: new Abstract: Constitutional AI is a method to oversee and control LLMs based on a set of rules written in natural language. …

Rushil Thareja, Gautam Gupta, Francesco Pinto, Nils Lukas
조회수 42회
Academic · 약 1분

MOSAIC: Composable Safety Alignment with Modular Control Tokens

arXiv:2603.16210v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is commonly implemented as a single static policy embedded in model parameters. However, …

Jingyu Peng, Hongyu Chen, Jiancheng Dong, Maolin Wang, Wenxi Li, Yuchen Li, Kai Zhang, Xiangyu Zhao
조회수 56회