作者
Chelsea Finn
Robotics / Machine Learning
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
机构 * Physical Intelligence
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Comments Project website: https://droid-dataset.github.io/
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Comments Project website: https://cot-vla.github.io/
Journal ref CVPR 2025
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
Comments ICRA 2024
Robotic Control via Embodied Chain-of-Thought Reasoning
Comments Project Website: https://embodied-cot.github.io. Updated funding information
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
FAST: Efficient Action Tokenization for Vision-Language-Action Models
Comments Website: https://www.pi.website/research/fast
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
Correct-N-Contrast: A Contrastive Approach for Improving Robustness to Spurious Correlations
Comments 38 pages, 14 figures. ICML 2022 Long Talk
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Comments 30 pages, 38th Conference on Neural Information Processing Systems (NeurIPS 2024)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
Generative Reward Models
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
Comments Project website: https://helpful-doggybot.github.io/
Calibrating Language Models with Adaptive Temperature Scaling
Comments EMNLP 2024
Disentangling Length from Quality in Direct Preference Optimization
OpenVLA: An Open-Source Vision-Language-Action Model
Comments Website: https://openvla.github.io/
Clarify: Improving Model Robustness With Natural Language Corrections
Comments UIST 2024. Interface code available at https://github.com/yoonholee/Clarify
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
Comments RLC 2024
Efficient Imitation Learning with Conservative World Models
Comments Oral presentation, L4DC 2024
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Comments COLM 2024
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
PERSONA: A Reproducible Testbed for Pluralistic Alignment
Surgical Robot Transformer (SRT): Imitation Learning for Surgical Tasks
Comments 8 pages
Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Comments 42 pages, 13 figures, 33 tables
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
Comments 27 pages