Zum Inhalt springen
  • Data Analytics and Machine Learning Group
  • TUM School of Computation, Information and Technology
  • Technische Universität München
Technische Universität München
  • Startseite
  • Team
    • Stephan Günnemann
    • Sirine Ayadi
    • Tim Beyer
    • Jonas Dornbusch
    • Eike Eberhard
    • Dominik Fuchsgruber
    • Nicholas Gao
    • Lukas Gosch
    • Filippo Guerranti
    • Leon Hetzel
    • Chengzhi Martin Hu
    • Niklas Kemper
    • Amine Ketata
    • Marcel Kollovieh
    • Arthur Kosmala
    • Aleksei Kuvshinov
    • Richard Leibrandt
    • Marten Lienen
    • David Lüdke
    • Aman Saxena
    • Sebastian Schmidt
    • Yan Scholten
    • Jan Schuchardt
    • Leo Schwinn
    • Johanna Sommer
    • Tim Tomov
    • Tom Wollschläger
    • Alumni
      • Simon Geisler
      • Anna-Kathrin Kopetzki
      • Amir Akbarnejad
      • Roberto Alonso
      • Bertrand Charpentier
      • Marin Bilos
      • Aleksandar Bojchevski
      • Johannes Gasteiger, né Klicpera
      • Maria Kaiser
      • Richard Kurle
      • Hao Lin
      • John Rachwan
      • Oleksandr Shchur
      • Armin Moin
      • Daniel Zügner
  • Lehre
    • Wintersemester 2026/27
      • Machine Learning
    • Sommersemester 2026
      • Machine Learning for Graphs and Sequential Data
      • Advanced Machine Learning: Deep Generative Models
      • Applied Machine Learning
    • Wintersemester 2025/26
      • Machine Learning
      • Robust Machine Learning
      • Seminar: Current Topics in Machine Learning
      • Seminar: Selected Topics in Machine Learning Research
    • Sommersemester 2025
      • Advanced Machine Learning: Deep Generative Models
      • Applied Machine Learning
      • Seminar: Selected Topics in Machine Learning Research
      • Seminar: Current Topics in Machine Learning
    • Wintersemester 2024/25
      • Machine Learning
      • Seminar: Selected Topics in Machine Learning Research
      • Seminar: Current Topics in Machine Learning
    • Sommersemester 2024
      • Machine Learning for Graphs and Sequential Data
      • Advanced Machine Learning: Deep Generative Models
      • Applied Machine Learning
      • Seminar: Selected Topics in Machine Learning Research
    • Wintersemester 2023/24
      • Machine Learning
      • Applied Machine Learning
      • Seminar: Selected Topics in Machine Learning Research
      • Seminar: Machine Learning for Sequential Decision Making
    • Sommersemester 2023
      • Machine Learning for Graphs and Sequential Data
      • Advanced Machine Learning: Deep Generative Models
      • Large-Scale Machine Learning
      • Seminar
    • Wintersemester 2022/23
      • Machine Learning
      • Large-Scale Machine Learning
      • Seminar
    • Sommersemester 2022
      • Machine Learning for Graphs and Sequential Data
      • Large-Scale Machine Learning
      • Seminar (Selected Topics)
      • Seminar (Time Series)
    • Wintersemester 2021/22
      • Machine Learning
      • Large-Scale Machine Learning
      • Seminar
    • Sommersemester 2021
      • Machine Learning for Graphs and Sequential Data
      • Large-Scale Machine Learning
      • Seminar
    • Wintersemester 2020/21
      • Machine Learning
      • Large-Scale Machine Learning
      • Seminar
    • Sommersemester 2020
      • Machine Learning for Graphs and Sequential Data
      • Large-Scale Machine Learning
      • Seminar
    • Wintersemester 2019/20
      • Machine Learning
      • Large-Scale Machine Learning
    • Sommersemester 2019
      • Mining Massive Datasets
      • Large-Scale Machine Learning
      • Oberseminar
    • Wintersemester 2018/19
      • Machine Learning
      • Large-Scale Machine Learning
      • Oberseminar
    • Sommersemester 2018
      • Mining Massive Datasets
      • Large-Scale Machine Learning
      • Oberseminar
    • Wintersemester 2017/18
      • Machine Learning
      • Oberseminar
    • Sommersemester 2017
      • Robust Data Mining Techniques
      • Efficient Inference and Large-Scale Machine Learning
      • Oberseminar
    • Wintersemester 2016/17
      • Mining Massive Datasets
    • Sommersemester 2016
      • Large-Scale Graph Analytics and Machine Learning
    • Wintersemester 2015/16
      • Mining Massive Datasets
    • Sommersemester 2015
      • Data Science in the Era of Big Data
    • Machine Learning Lab
  • Forschung
    • Robust Machine Learning
    • Machine Learning for Graphs/Networks
    • Machine Learning for Temporal and Dynamical Data
    • Bayesian (Deep) Learning / Uncertainty
    • Efficient ML
    • Code
  • Publikationen
  • Offene Stellen
    • FAQ
  • Abschlussarbeiten
  1. Startseite
  2. Forschung

A Unified Approach Towards Active Learning and Out-of-Distribution Detection

by Sebastian Schmidt, Leonard Schenk, Leo Schwinn, and Stephan Günnemann.
Transactions on Machine Learning Research (TMLR), 2025.

Paper |  Code

Abstract:

In real-world applications of deep learning models, active learning (AL) strategies are essential for identifying label candidates from vast amounts of unlabeled data. In this context, robust out-of-distribution (OOD) detection mechanisms are crucial for handling data outside the target distribution during the application’s operation. Usually, these problems have been addressed separately. In this work, we introduce SISOM as a unified solution designed explicitly for AL and OOD detection. By combining feature space-based and uncertainty-based metrics, SISOM leverages the strengths of the currently independent tasks to solve both effectively without requiring specific training schemes. We conducted extensive experiments showing the problems arising when migrating between the two tasks. In our experiments, SISOM underlined its effectiveness by achieving first place in one of the commonly used OpenOOD benchmark settings and top-3 places in the remaining two for near-OOD data. In AL, SISOM 1 delivers top performance in common image benchmarks.

Bridging the Gap in Real-World Deep Learning

Deploying large-scale deep learning models in real-world applications, such as autonomous driving or robotic perception, introduces two primary data-centric challenges. First, models require vast amounts of annotated data to handle uncontrolled environments, a bottleneck traditionally mitigated by Active Learning (AL) strategies that guide the selection of unlabeled candidates. Second, even extensively trained models can behave unpredictably when they encounter out-of-distribution (OOD) data that deviates significantly from their training set.

Historically, AL and OOD detection have been treated as entirely independent tasks requiring separate methodological components and training overheads. However, as outlined in our work there is a distinct ambiguity and overlap between unlabeled data that is highly informative for AL and near-OOD data. Both types of samples typically reside in the complex boundary regions between known data clusters. By recognizing this entanglement, we can simplify the application life cycle and eliminate the need for separate, often conflicting, task designs

Simultaneous Informative Sampling and Outlier Mining

To address these challenges holistically, we introduce Simultaneous Informative Sampling and Outlier Mining (SISOM), a unified model that seamlessly performs both AL and OOD detection. By combining feature-space-based distances with uncertainty-based metrics, SISOM achieves state-of-the-art performance across both domains without relying on any specific training schemes.

The SISOM architecture targets the unexplored boundary regions between data clusters using several key mechanisms:

  • Expanded Coverage: Instead of relying on a single layer, SISOM defines its feature-space representation by concatenating latent spaces from multiple network layers to maximize information gain.
  • Feature Enhancement: The model increases class separation by applying a saliency weighting to individual neurons. This is achieved using the gradient of the Kullback-Leibler (KL) divergence between a uniform distribution and the model's softmax output.
  • Distance Ratio Metric: SISOM identifies vital samples by computing the ratio of inner-class to outer-class distances. It compares the minimum distance to a sample from the predicted class with the minimum distance to a sample from a different class.
  • Self-Deciding Fusion (SISOMe): To account for datasets in which feature spaces are difficult to separate, the framework includes a self-balancing analysis of feature spaces. This component optimally weights the SISOM diversity metric against an uncertainty-based energy score.
  • Sigmoid Steepness Optimization: The framework enables post-training refinement of feature-space representations by optimizing the steepness parameters of the sigmoid function applied across different layers.
To top

Informatik 26 - Data Analytics and Machine Learning


Prof. Dr. Stephan Günnemann

Technische Universität München
TUM School of Computation, Information and Technology
Department of Computer Science
Boltzmannstr. 3
85748 Garching 

Sekretariat:
Raum 00.11.057
Tel.: +49 89 289-17256
Fax: +49 89 289-17257

  • Datenschutz
  • Impressum
  • Barrierefreiheit