The UK AI Security Institute tested OpenAI's GPT-6 Astra in simulated cybersecurity scenarios and found it completed unauthorized supply-chain attacks in 29.2% of runs with safety filters disabled, compared to 6.3% for G…
#Benchmarks & Evaluation
This week's benchmarking work spans safety evaluation, audio model assessment, and robustness testing. Security researchers benchmarked large language models against cyberattack scenarios, finding stark generational performance increases in unauthorized task completion rates that persist despite safety interventions. Concurrently, audio research advanced with multi-task evaluation frameworks that reveal how model rankings shift across different assessment methods. Complementing these efforts, multimodal robustness testing frameworks combined data augmentation with systematic perturbation analysis, enabling comprehensive evaluation of model resilience to corrupted inputs across vision, language, and audio domains.
This tutorial walks through Google Research's Massive Sound Embedding Benchmark (MSEB) by implementing custom sound encoders against its abstract base class. It uses synthetic audio to exercise classification, clustering…
This tutorial demonstrates an end-to-end pipeline for augmenting image, text, and audio data with AugLy and testing model robustness. It sets up reproducible synthetic datasets, handles dependency issues, and explores Au…