Making Sense of Touch from the Child's View for Contrastive Learning

ICDL 2026
Max Whitton1,  Zecheng Wang*1,  Puchen Liu*1,  Quang Tuan Truong1,  Shengao Wang1,  Manaswi Yadamreddy1,  Oktay Ozel1,  Visista Jayanti1,  Saniya Sekhon1,  Hanna Samuel Tadesse1,  Lawrence Miao1,  Junjie Wang1,  Jiasen Lu2,  Chen Yu3,  Boqing Gong1
1 Boston University  |  2 Apple  |  3 The University of Texas at Austin
* Equal contribution
Boston University
Apple
UT Austin

Abstract

Is the sense of touch a mechanism for human babies' learning of visual concepts? If so, can we quantify its importance, and to what extent do babies rely on their sense of touch for visual learning? To approach these questions in a principled way, we propose a structured coding system for baby-centric touch events, yielding a dataset of 264k two-second clips of touch events coded according to this system. Using this dataset, we pretrain developmentally grounded models that reveal promising insights into the nature of baby learning from touch.

Data Pipeline

Three-stage pipeline to code a large-scale dataset of baby-centric touch events.
Data pipeline

Dataset Statistics

Distributions of the coded touch [verb, noun, body part, additional conditions] over our touch annotations.
Touch annotation distributions

Data Examples

Examples of [frame, speech transcript, touch ID] triplets (top) and data pairs missing either speech or touch (bottom).
Data examples

Model Architecture

Vision, touch, and text encoders are trained so that representations of the same scene are pushed together while representations of different scenes are pushed apart.
Model architecture

Evaluation Format

Labeled-S: Given a query concept, select the matching frame from a set of candidates.
Labeled-S evaluation format
Picture Vocabulary: Given a spoken word, identify the correct image from four options.
Picture Vocabulary evaluation format

Qualitative Results

Representative pretraining examples (left), and corresponding attention maps of our model (middle) and a baseline model (right).
Attention maps

Quantitative Results

Linear probe and zero-shot accuracy across three tasks. Blue = touch-enhanced pretraining, red = baseline.
All quantitative results