MOSS-VL · an open 11B video model trained to decide, after every frame, whether to speak or stay silent