Samuel Clarke

Located in Palo Alto, CA
spclarke [at] alumni [dot] stanford [dot] edu

I work at Tesla AI, on the Foundation Models team. I finished my PhD in Computer Science at Stanford University in 2025, advised by Professor Jiajun Wu. Previously, I was a Graduate Research Assistant in Carnegie Mellon's Robotics Institute, co-advised by Professors Chris Atkeson and Oliver Kroemer. My most recent research has been in the realm of learning for manipulation, especially how robots can use sound and tactile sensing during manipulation.

Samuel Clarke

Research

DexSkin: High-Coverage Conformable Robotic Skin for Learning Contact-Rich Manipulation

Suzannah Wistreich*, Baiyu Shi*, Stephen Tian*, Samuel Clarke, Michael Nath, Chengyi Xu, Zhenan Bao, and Jiajun Wu

CoRL 2025

We introduce a soft, capacitive electronic skin that conforms to curved robotic surfaces and provides localized tactile sensing with high coverage. We show that it enables manipulation tasks requiring sensing across the whole finger, to support learning contact-rich manipulation in the real world.

X-Capture: An Open-Source Portable Device for Multi-Sensory Learning

Samuel Clarke, Suzannah Wistreich, Yanjie Ze, and Jiajun Wu

ICCV 2025

We design an open-source, portable, and cost-effective device for real-world multi-sensory data collection, capturing correlated RGBD images, tactile readings, and impact audio for under $1,000 in parts. We use this device to curate a dataset of 3,000 points on 500 everyday objects from diverse real-world environments, and show that both the quantity and sensory breadth of this data help in pretraining and fine-tuning multi-modal representations for object-centric tasks.

Hearing Anything Anywhere

Mason Wang*, Ryosuke Sawata*, Samuel Clarke, Ruohan Gao, Shangzhe Wu, and Jiajun Wu

CVPR 2024

We collect a dataset of real Room Impulse Responses (RIRs) from four rooms, then introduce a new Novel Viewpoint Acoustic Synthesis method based on differentiable audio rendering. Our method uses physics-based biases to achieve practical sample efficiency, only requiring RIRs to be measured at roughly 12 distinct locations in the room.

SoundCam: A Dataset for Finding Humans Using Room Acoustics

Mason Wang*, Samuel Clarke*, Jui-Hsien Wang, Ruohan Gao, and Jiajun Wu

NeurIPS Datasets & Benchmarks 2023

We record thousands of room impulse responses and music clips in different real rooms with humans standing in different positions in the room. Learning-based models can use these minute differences in the room's acoustics to track, identify, or detect humans in the room. Our data can be used to develop more robust and sample-efficient methods, with applications in home assistants, security, and robotics.

RoboCook: Long-Horizon Elasto-Plastic Object Manipulation with Diverse Tools

Haochen Shi*, Huazhe Xu*, Samuel Clarke, Yunzhu Li, and Jiajun Wu

CoRL 2023 Best Systems Paper

We introduce a framework for accomplishing long-horizon tasks in soft-body manipulation and show that it can learn to make dumplings with a variety of tools, with very little training data for each tool. We also show that our framework can learn to use tools to achieve other tasks in soft-body manipulation, such as shaping dough into target shapes, autonomously selecting tools for each step of the task.

RealImpact: A Dataset of Impact Sound Fields for Real Objects

Samuel Clarke, Ruohan Gao, Mason Wang, Mark Rau, Julia Xu, Jui-Hsien Wang, Doug James, and Jiajun Wu

CVPR 2023 Highlight

We collect 150,000 annotated recordings of impacts of 50 everyday objects, recorded from 600 distinct microphone locations. We show how our data can be used to tune and validate acoustic simulations, or used directly in interesting downstream audio and audiovisual tasks.

DiffImpact: Differentiable Rendering and Identification of Impact Sounds

Samuel Clarke, Negin Heravi, Mark Rau, Ruohan Gao, Jiajun Wu, Doug James, and Jeannette Bohg

CoRL 2021 Oral Presentation

Differentiable physics-based models provide a useful bias for learning from impact sounds to solve both forward and backward problems on impact audio. We show we can both infer models from data in the wild, and then use these models to perform source separation better than generic learning-based alternatives.

Robot Learning for Manipulation of Granular Materials Using Vision and Sound

Samuel Clarke (in collaboration with Travers Rhodes and advisors Christopher G. Atkeson and Oliver Kroemer)

CMU Masters in Robotics Thesis 2019

Deep learning-based data-driven models can both predict the effects of a scooping operation on a granular material using vision and can learn to use audio for feedback on scooping and pouring granular materials.

Learning Audio Feedback for Estimating Amount and Flow of Granular Material

Samuel Clarke, Travers Rhodes, Christopher G. Atkeson, and Oliver Kroemer

CoRL 2018

Deep learning-based data-driven models can accurately predict the amount of granular materials a robot pours or shakes, based only on audio recordings. With machine learning, we can use recordings from a $3 microphone to outperform the measurement resolution of a $3,000 wrist force-torque sensor.

NaturalMotion: Exploring Gesture Controls for Visualizing Time-Evolving Graphs

Samuel Clarke, Nathan Dass, and Polo Chau

VIS 2016

We took the Matrix Cube, a tool for visualizing time-evolving graphs, and developed a new 3D interface controlled by hand gestures. Gestures were captured with a Leap Motion device.

LatentGesture: Active User Authentication through Background Touch Analysis

Premkumar Saravanan, Samuel Clarke, Polo Chau, and Hongyuan Zha

Chinese CHI 2014

With data collected from a user's brief interactions with common UI elements, such as checkboxes and sliders, machine learning models can uniquely identify the user. Such a system could authenticate a mobile device's user continuously and seamlessly, without many of the vulnerabilities common to traditional authentication methods.

Projects

CPR+ Mask

Finalist for Georgia Tech's 2017 Inventure Prize

We designed a mask to walk an untrained rescuer through performing standard-of-care CPR to a cardiac arrest victim, in real time. The mask is equipped with sensors for monitoring the state of the victim and the quality of the CPR and uses a speaker and LEDs to give instructions and cues to the rescuer.

Experience

Senior Machine Learning Engineer

August 2025 – Present

Tesla AI

  • On Foundation Models team.

Graduate Research Assistant

Sep 2019 – Aug 2025

Stanford University (advised by Jiajun Wu)

  • Developed multi-modal embedding for interactive robot perception of objects (see Research).
  • Developed framework for learning robot manipulation with novel tactile sensor (see Research).
  • Developed learning-based methods for deriving object modal sound models from a dataset of real recordings of object impact sounds (see Research).
  • Developed novel automated system for collecting large datasets of real sounds from everyday objects for acoustic model benchmarking (see Research).
  • Developed methods for actively tracking humans in rooms using audio (see Research).

Research Intern

Jun 2022 – Nov 2022

Adobe Research

  • Synthesized datasets and developed learning-based methods for tracking object motion using sound.

Machine Learning Intern

May 2019 – Aug 2019

Autoroboto at Google Brain

  • Developed deep learning-based methods to extract feedback and detect anomalies from audio, with applications in industrial automation.

Graduate Research Assistant

Nov 2017 – May 2019

Carnegie Mellon University (co-advised by Chris Atkeson and Oliver Kroemer)

  • Developed deep learning-based audio feedback framework for estimating masses for robotic pouring and scooping granular materials (see Research).
  • Developed deep learning-based data-driven framework for predicting effects of robotic scooping granular materials (see Research).
  • Designed and manufactured numerous parts for lab projects using SolidWorks.
  • Designed and trained many learning models using TensorFlow.

Mechanical Engineering Intern

May 2016 – Aug 2016

Google

  • Led multi-team design and implementation of an automated vision-based scanning device to ensure a secure exit of hard drives from data centers.
  • Developed APIs for automation machines to query desired data from sources on production networks with Go, Python, and C++ communicating through JSON.

Software Engineering Intern

Apr 2015 – Aug 2015

Google

  • Built Arduino-controlled motorized linear slide to automate Android device tests.
  • Developed Go adapter for communicating with Arduino devices over serial.
  • Developed tools for compiling Arduino code within company codebase.
  • Wrote automated communications tests for Android devices, using Go and Android Java.

Propulsion Data Science Intern

Jan 2015 – Apr 2015

SpaceX

  • Collaborated in development of an automated system to detect anomalies in telemetry sensor data from rocket engine tests, targeting a 12x reduction in human review time.
  • Transferred legacy academic code modelling the Merlin 1D rocket engine from Fortran to Python to optimize performance of a new iteration of the engine.

Undergraduate Researcher

Jan 2013 – Dec 2016

Georgia Tech (advised by Polo Chau)

  • Developed gesture-based 3D data visualization in an Oculus Rift virtual reality environment, with demo developed in the Unity game engine (see Research).
  • Developed method for behavioral authentication on Android featured in popular press (Engadget, Gizmodo, Yahoo, many more) (see Research).

Engineering Practicum Intern

May 2014 – Aug 2014

Google

  • Developed backend RPC infrastructure for serving feature search requests to the search bar in Google Maps Engine viewer.
  • Mechanical 20% project with Glass: studied repeatability of 4-axis automated manufacturing robots through designing tests and analyzing results in MATLAB.

Freshman Engineering Practicum Intern

May 2013 – Aug 2013

Google

  • Developed automated user interface tests for Google Play on Android.
  • Developed App Engine project in Python to organize and monitor test results.
  • 20% project with Android hardware team: designed and prototyped testing tool and façade case for unannounced products in Creo Parametric, performed mechanical tests.

Contractor

Apr 2012 – Jul 2012

Salesforce Marketing Cloud

  • Prototyped HBase infrastructure and an application to efficiently move billions of records from SQL Server to Hadoop and HBase, as a feasibility study.
  • Worked with C#, Pig, Java, and Thrift.

Intern

Sep 2011 – Apr 2012

ChaCha Search

  • Data mining and analysis tasks using Hadoop, Pig, Tableau, and Excel.
  • Developed in-house click fraud detection algorithm which rivaled proprietary options.

Education

Stanford University

Sep 2019 – Aug 2025

Ph.D., Computer Science · GPA 4.21/4.3

Carnegie Mellon University

Aug 2017 – May 2019

M.S., Robotics · GPA 4.17/4.33

Georgia Institute of Technology

Aug 2012 – May 2016

B.S., Computer Science and Mechanical Engineering · GPA 4.0/4.0