Saubhagya Pandey

← Door A · Engineering & AI

FIG. A.5 · Research · DCASE 2025 Challenge Task 6

Language-Based Audio Retrieval

Find a sound from a sentence: “a dog barking while rain hits a tin roof.” This is my DCASE 2025 Task 6 submission, built under the supervision of Prof. Vipul Arora, IIT Kanpur, on top of the challenge’s dual-encoder baseline. The contribution is in the training recipe — which pairs to treat as positives, which negatives are worth learning from, and when to turn each mechanism on.

30.01

mAP@10 · Clotho · official

59.35

R@10 · Clotho · official

0.75

soft-positive threshold

0.15

hard-negative margin

Leaderboard figures from the official DCASE 2025 Task 6 evaluation of this submission. The repository isolates each component’s contribution via ablations; only the official metrics are reproduced here.

THE TRAINING RECIPE — STAGED, NOT SIMULTANEOUS

STEP 01

Multi-positive learning

Semantically similar samples are treated as soft positives with weighted loss, at a 0.75 similarity threshold — one caption isn't the only right match for a clip.

STEP 02

Hard-negative mining

Top-5 hardest negatives per anchor (margin 0.15) concentrate training on the confusing cases instead of the easy ones.

STEP 03

Progressive curriculum

Techniques activate in stages by epoch rather than all at once, letting the shared embedding space form before the hard cases arrive.

TWO-STAGE INFERENCE

PASS 01

Bi-encoder retrieval

PaSST audio tower and RoBERTa-large text tower, trained into a shared embedding space with InfoNCE, retrieve the top-K candidates fast.

PASS 02

Compositional rerank

Multimodal transformer adapters rerank the shortlist compositionally, where cross-modal attention can see parts of the caption the towers averaged away.

WHY THIS MATTERED

Retrieval benchmarks reward whatever trick you bolt on; the discipline here was deciding what earns its place. Each mechanism — soft positives, hard negatives, the curriculum, the rerank stage — had to justify itself against the baseline before staying in the submission. The result is a ranked leaderboard entry rather than a collection of plausible ideas.

Prev: FIG. A.4 · Slack Data CopilotNext: FIG. A.6 · Hindi Sahitya Sabha Website