721 lines
117 KiB
JSON
721 lines
117 KiB
JSON
{
|
||
"candidates": [
|
||
{
|
||
"title": "Visuomotor policy learning via action diffusion - ACM Digital Library",
|
||
"snippet": "Diffusion policy: : Visuomotor policy learning via action diffusion: International Journal of Robotics Research: Vol 4\n[...]\n, No 1\n[...]\nother-periodical;requested\n[...]\n:string:\n[...]\nPublication Websites;subPage:\n[...]\n:Basic Abstract\n[...]\n;page:string:Article/Chapter View;ctype:string:Journal Content;group\n[...]\n:acm-\n[...]\ntype>other-periodical;website:website:\n[...]\n-site;\n[...]\n:issue:\n[...]\n\\:10.5\n[...]\n55/rbrs.20\n[...]\n.44.issue-10-11;csubtype:string:Periodical;taxonomy:taxonomy:acm-pubtype;pageGroup:string:Publication Pages\"> skip to main content\n\n \n\n \n\nContents\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 15 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion mode",
|
||
"source_url": "https://dl.acm.org/doi/10.1177/02783649241273668",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://dl.acm.org/doi/10.1177/02783649241273668",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Visuomotor Policy Learning via Action Diffusion - arXiv",
|
||
"snippet": "# Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 15 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details is available diffusion-policy.cs.columbia.edu\n[...]\nIn this work, we s",
|
||
"source_url": "https://arxiv.org/abs/2303.04137",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2303.04137",
|
||
"_exa_published_date": "2023-03-07T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2303.04137v5] Diffusion Policy: Visuomotor Policy Learning via Action Diffusion",
|
||
"snippet": "[2303.04137v5] Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\n# Title:Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\n> Abstract:This paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot's visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models.",
|
||
"source_url": "http://arxiv.org/abs/2303.04137v5",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "http://arxiv.org/abs/2303.04137v5",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[2303.04137] Diffusion Policy - ar5iv - arXiv",
|
||
"snippet": "# Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details will be publicly available.\n[...]\nIn this work, we seek to address thi",
|
||
"source_url": "https://ar5iv.labs.arxiv.org/html/2303.04137",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://ar5iv.labs.arxiv.org/html/2303.04137",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[2006.11239v2] Denoising Diffusion Probabilistic Models",
|
||
"snippet": "[2006.11239\n[...]\n] Denoising Diffusion Probabilistic Models\n[...]\n# Title:Denoising Diffusion Probabilistic Models\n[...]\nAuthors: Jonathan Ho, Ajay Jain, Pieter Abbeel\n[...]\n> Abstract:We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our implementation is available at this https URL\n\n \n\nhttps://doi.org/10.48550/arXiv.2006.11239\n\n \n\narXiv-issued DOI via DataCite",
|
||
"source_url": "https://arxiv.org/abs/2006.11239v2",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2006.11239v2",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[2006.11239] Denoising Diffusion Probabilistic Models - arXiv",
|
||
"snippet": "[2006.11239] Denoising Diffusion Probabilistic Models\n[...]\n# Denoising Diffusion Probabilistic Models\n[...]\nJonathan Ho UC Berkeley jonathanho@berkeley.edu &Ajay Jain UC Berkeley ajayj@berkeley.edu &Pieter Abbeel UC Berkeley pabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our implementation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [53]. A diffusion probabilistic model (which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using variational inference to produce samples matching the data after finite time. Transitions of this chain are learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the data in the opposite direction of s",
|
||
"source_url": "https://arxiv.org/abs/2006.11239",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2006.11239",
|
||
"_exa_published_date": "2020-06-19T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Denoising Diffusion Probabilistic Models",
|
||
"snippet": "Denoising Diffusion Probabilistic Models \n\nAuthorFeedback Bibtex MetaReview Paper Review Supplemental\n\n## Abstract\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN.\n\n \n\nDo not remove: This comment is monitored to verify that the site is working properly",
|
||
"source_url": "https://papers.nips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://papers.nips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "",
|
||
"snippet": "Denoising Diffusion Probabilistic Models\n[...]\njonathanho@berkeley.edu\n[...]\nAjay Jain\n[...]\najayj@berkeley.edu\n[...]\nPieter Abbeel\n[...]\npabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models,\n[...]\na class of latent variable models inspired by considerations from nonequilibrium\n[...]\nthermodynamics. Our best results are obtained by training on a weighted variational\n[...]\nbound designed according to a novel connection between diffusion probabilistic\n[...]\nmodels and denoising score matching with Langevin dynamics, and our models nat\u0002urally admit a progressive lossy decompression scheme that can be interpreted as a\n[...]\ngeneralization of autoregressive decoding. On the unconditional CIFAR10 dataset,\n[...]\nwe obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On\n[...]\n256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our imple\u0002mentation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [50]. A diffusion probabilistic model\n[...]\n(which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using\n[...]\nvariational inference to produce samples matching the data after finite time. Transitions of this chain\n[...]\nare learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the\n[...]\ndata in the opposite direction of sampling until signal",
|
||
"source_url": "https://arxiv.org/pdf/2006.11239v1",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy",
|
||
"rw.mean_flow",
|
||
"rw.vla",
|
||
"rw.attnres",
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2006.11239v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Denoising Diffusion Probabilistic Models - arXiv",
|
||
"snippet": "Denoising Diffusion Probabilistic Models\n[...]\njonathanho@berkeley.edu\n[...]\nAjay Jain\n[...]\najayj@berkeley.edu\n[...]\nPieter Abbeel\n[...]\npabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models,\n[...]\na class of latent variable models inspired by considerations from nonequilibrium\n[...]\nthermodynamics. Our best results are obtained by training on a weighted variational\n[...]\nbound designed according to a novel connection between diffusion probabilistic\n[...]\nmodels and denoising score matching with Langevin dynamics, and our models nat\u0002urally admit a progressive lossy decompression scheme that can be interpreted as a\n[...]\ngeneralization of autoregressive decoding. On the unconditional CIFAR10 dataset,\n[...]\nwe obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On\n[...]\n256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our imple\u0002mentation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [53]. A diffusion probabilistic model\n[...]\n(which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using\n[...]\nvariational inference to produce samples matching the data after finite time. Transitions of this chain\n[...]\nare learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the\n[...]\ndata in the opposite direction of sampling until signal",
|
||
"source_url": "https://arxiv.org/pdf/2006.11239",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2006.11239",
|
||
"_exa_published_date": "2020-12-16T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2212.09748] Scalable Diffusion Models with Transformers - arXiv",
|
||
"snippet": "William Peebles* UC Berkeley Saining Xie New York University\n[...]\nWe explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops—through increased transformer depth/width or increased number of input tokens—consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512 $\\times$ 512 and 256 $\\times$ 256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.\n[...]\n, or Di\n[...]\nfor short. Di\n[...]\nadhere to the best practices of Vision Transformers (ViTs) [10], which have been shown to scale more effectively for visual recognition than\n[...]\nconvolutional networks (e.g., Res\n[...]\n[15]).\n[...]\nMore specifically, we study the scaling behavior of transformers with respect to network complexity vs. sample quality. We show that by constructing and benchmarking the DiT design space under the Latent Diffusion Models (LDMs) [48] framework, where diffusion models are trained within a VAE’s latent space, we can successfully replace the U-Net backbone with a transformer. We further show that DiTs are scalable architectures for diff",
|
||
"source_url": "https://arxiv.org/abs/2212.09748",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2212.09748",
|
||
"_exa_published_date": "2022-12-19T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore",
|
||
"snippet": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore\n\n \n\n \n\n \n\n### IEEE Account",
|
||
"source_url": "https://ieeexplore.ieee.org/document/10377858/",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://ieeexplore.ieee.org/document/10377858/",
|
||
"_exa_published_date": "2025-05-14T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2212.09748v1] Scalable Diffusion Models with Transformers",
|
||
"snippet": "[2212.09748v1] Scalable Diffusion Models with Transformers\n[...]\n# Title:Scalable Diffusion Models with Transformers\n[...]\nAuthors: William Peebles, Saining Xie\n[...]\n> Abstract:We explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops -- through increased transformer depth/width or increased number of input tokens -- consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512x512 and 256x256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.",
|
||
"source_url": "http://arxiv.org/abs/2212.09748v1",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "http://arxiv.org/abs/2212.09748v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Scalable Diffusion Models with Transformers | Semantic Scholar",
|
||
"snippet": "[PDF] Scalable Diffusion Models with Transformers | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Peebles2022ScalableDM,\n title={Scalable Diffusion Models with Transformers},\n author={William S. Peebles and Saining Xie},\n journal={2023 IEEE/CVF International Conference on Computer Vision (ICCV)},\n year={2022},\n pages={4172-4182},\n url={https://api.semanticscholar.org/CorpusID:254854389}\n}\n```",
|
||
"source_url": "https://www.semanticscholar.org/reader/736973165f98105fec3729b7db414ae4d80fcbeb",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://www.semanticscholar.org/reader/736973165f98105fec3729b7db414ae4d80fcbeb",
|
||
"_exa_published_date": "2022-12-19T14:39:50.000Z"
|
||
},
|
||
{
|
||
"title": "Scalable Diffusion Models with Transformers - IEEE Xplore",
|
||
"snippet": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore\n**\n\n### IEEE Account\n* Change Username/Password\n* Update Address\n### Purchase Details\n* Payment Options\n* Order History\n* View Purchased Documents\n### Profile Information\n* Communications Preferences\n* Profession and Education\n* Technical Interests\n### Need Help?\n* **US & Canada:**+1 800 678 4333\n* **Worldwide:**+1 732 981 0060\n* Contact & Support\n* About IEEE*Xplore*\n* Contact Us\n* Help\n* Accessibility\n* Terms of Use\n* Nondiscrimination Policy\n* Sitemap\n* Privacy & Opting Out of Cookies\nA not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.\n© Copyright 2025 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.\n**",
|
||
"source_url": "https://ieeexplore.ieee.org/iel7/10376473/10376477/10377858.pdf",
|
||
"discovered_for": [
|
||
"rw.diffusion_policy"
|
||
],
|
||
"_exa_id": "https://ieeexplore.ieee.org/iel7/10376473/10376477/10377858.pdf",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Flow Matching for Generative Modeling - arXiv",
|
||
"snippet": "FLOW MATCHING FOR GENERATIVE MODELING\n[...]\nYaron Lipman1,2 Ricky T. Q. Chen1 Heli Ben-Hamu2 Maximilian Nickel1 Matt Le1\n[...]\nWe introduce a new paradigm for generative modeling built on Continuous\n[...]\n(CNFs), allowing us to train CNFs at unprecedented scale.\n[...]\nSpecifically, we present the notion of Flow Matching (FM), a simulation-free\n[...]\napproach for training CNFs based on regressing vector fields of fixed conditional\n[...]\nprobability paths. Flow Matching is compatible with a general family of Gaussian\n[...]\nprobability paths for transforming between noise and data samples—which\n[...]\nsubsumes existing diffusion paths as specific instances. Interestingly, we find\n[...]\nthat employing FM with diffusion paths results in a more robust and stable\n[...]\nalternative for training diffusion models. Furthermore, Flow Matching opens\n[...]\nthe door to training CNFs with other, non-diffusion probability paths. An\n[...]\ninstance of particular interest is using Optimal Transport (OT) displacement\n[...]\ninterpolation to define the conditional probability paths. These paths are more\n[...]\nefficient than diffusion paths, provide faster training and sampling, and result in\n[...]\nbetter generalization. Training CNFs using Flow Matching on ImageNet leads\n[...]\nto consistently better performance than alternative diffusion-based methods in\n[...]\nterms of both likelihood and sample quality, and allows fast and reliable sample\n[...]\ngeneration using off-the-shelf numerical ODE solvers.\n",
|
||
"source_url": "https://arxiv.org/pdf/2210.02747",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2210.02747",
|
||
"_exa_published_date": "2023-02-08T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2210.02747] Flow Matching for Generative Modeling - arXiv",
|
||
"snippet": "[221\n[...]\n] Flow Matching for Generative Modeling\n[...]\n# Flow Matching for Generative Modeling\n[...]\nYaron Lipman1,2 Ricky T. Q. Chen1 Heli Ben-Hamu2 Maximilian Nickel1 Matt Le1 1Meta AI (FAIR) 2Weizmann Institute of Science\n[...]\nWe introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths. Flow Matching is compatible with a general family of Gaussian probability paths for transforming between noise and data samples—which subsumes existing diffusion paths as specific instances. Interestingly, we find that employing FM with diffusion paths results in a more robust and stable alternative for training diffusion models. Furthermore, Flow Matching opens the door to training CNFs with other, non-diffusion probability paths. An instance of particular interest is using Optimal Transport (OT) displacement interpolation to define the conditional probability paths. These paths are more efficient than diffusion paths, provide faster training and sampling, and result in better generalization. Training CNFs using Flow Matching on ImageNet leads to consistently better performance than alternative diffusion-based methods in terms of both likelihood and sample quality, and allows fast and reliable sample generation using off-the-s",
|
||
"source_url": "https://arxiv.org/abs/2210.02747",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2210.02747",
|
||
"_exa_published_date": "2022-10-06T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] Flow Matching for Generative Modeling | Semantic Scholar",
|
||
"snippet": "```\n@article{Lipman2022FlowMF,\n title={Flow Matching for Generative Modeling},\n author={Yaron Lipman and Ricky T. Q. Chen and Heli Ben-Hamu and Maximilian Nickel and Matt Le},\n journal={ArXiv},\n year={2022},\n volume={abs/2210.02747},\n url={https://api.semanticscholar.org/CorpusID:252734897}\n}\n```",
|
||
"source_url": "https://www.semanticscholar.org/reader/af68f10ab5078bfc519caae377c90ee6d9c504e9",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://www.semanticscholar.org/reader/af68f10ab5078bfc519caae377c90ee6d9c504e9",
|
||
"_exa_published_date": "2022-10-06T03:39:18.000Z"
|
||
},
|
||
{
|
||
"title": "Flow Matching for Generative Modeling - OpenReview",
|
||
"snippet": "Flow Matching for Generative Modeling | OpenReview\n\n## Flow Matching for Generative Modeling\n\nICLR 2023 notable top 25%Readers: Everyone\n\nKeywords: continuous normalizing flows, generative models\n\nAbstract: We introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths. Flow Matching is compatible with a general family of Gaussian probability paths for transforming between noise and data samples---which subsumes existing diffusion paths as specific instances. Interestingly, we find that employing FM with diffusion paths results in a more robust and stable alternative for training diffusion models. Furthermore, Flow Matching opens the door to training CNFs with other, non-diffusion probability paths. An instance of particular interest is using Optimal Transport (OT) displacement interpolation to define the conditional probability paths. These paths are more efficient than diffusion paths, provide faster training and sampling, and result in better generalization. Training CNFs using Flow Matching on ImageNet leads to consistently better performance than alternative diffusion-based methods in terms of both likelihood and sample quality, and allows fast and reliable sample generation using off-the-shelf numerical ODE solve",
|
||
"source_url": "https://openreview.net/forum?id=PqvMRDCJT9t",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=PqvMRDCJT9t",
|
||
"_exa_published_date": "2022-09-29T08:12:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2505.13447] Mean Flows for One-step Generative Modeling - arXiv",
|
||
"snippet": "# Mean Flows for One-step Generative Modeling\n[...]\nZhengyang Geng1 Mingyang Deng2 Xingjian Bai2 J. Zico Kolter1 Kaiming He2 1CMU 2MIT Work partly done when visiting MIT.\n[...]\nWe propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n[...]\nIn this work, we propose a principled and effective framework, termed MeanFlow, for one-step generation. The core idea is to introduce a new ground-truth field representing the average velocity, in contrast to the instantaneous velocity typically modeled in Flow Matching. Average velocity is defined as the ratio of displacement to a time interval, with displacement give",
|
||
"source_url": "https://arxiv.org/abs/2505.13447",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2505.13447",
|
||
"_exa_published_date": "2025-05-19T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Mean Flows for One-step Generative Modeling",
|
||
"snippet": "# Mean Flows for One-step Generative Modeling\n[...]\nZhengyang Geng1 Mingyang Deng2 Xingjian Bai2 J. Zico Kolter1 Kaiming He2 1CMU 2MIT Work partly done when visiting MIT.\n[...]\nWe propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n[...]\nIn this work, we propose a principled and effective framework, termed MeanFlow, for one-step generation. The core idea is to introduce a new ground-truth field representing the average velocity, in contrast to the instantaneous velocity typically modeled in Flow Matching. Average velocity is defined as the ratio of displacement to a time interval, with displacement give",
|
||
"source_url": "https://arxiv.org/html/2505.13447",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2505.13447",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Mean Flows for One-step Generative Modeling - OpenReview",
|
||
"snippet": "Mean Flows for One-step Generative Modeling | OpenReview\n\n## Mean Flows for One-step Generative Modeling\n\n### Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, Kaiming He\n\nNeurIPS 2025 oralEveryone Revisions BibTeX CC BY 4.0\n\nKeywords: Generative Models\n\nAbstract: We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the \\textit{MeanFlow} model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256$\\times$256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n\nPrimary Area: Deep learning (e.g., architectures, generative models, optimization for deep networks, foundation models, LLMs)\n\nSubmission Number: 754\n\nLoading",
|
||
"source_url": "https://openreview.net/forum?id=uWj4s7rMnR",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=uWj4s7rMnR",
|
||
"_exa_published_date": "2025-10-29T14:53:17.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] Mean Flows for One-step Generative Modeling | Semantic Scholar",
|
||
"snippet": "[PDF] Mean Flows for One-step Generative Modeling | Semantic Scholar\n[...]\n```\n@article{Geng2025MeanFF,\n title={Mean Flows for One-step Generative Modeling},\n author={Zhengyang Geng and Mingyang Deng and Xingjian Bai and J. Zico Kolter and Kaiming He},\n journal={ArXiv},\n year={2025},\n volume={abs/2505.13447},\n url={https://api.semanticscholar.org/CorpusID:278769814}\n}\n```",
|
||
"source_url": "https://www.semanticscholar.org/reader/19df654b0d0f634a451564346a09af8bd348dac0",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://www.semanticscholar.org/reader/19df654b0d0f634a451564346a09af8bd348dac0",
|
||
"_exa_published_date": "2025-05-19T06:36:24.000Z"
|
||
},
|
||
{
|
||
"title": "Improved Mean Flows: On the Challenges of Fastforward Generative ...",
|
||
"snippet": "# Improved Mean Flows: On the Challenges of Fastforward Generative Models\n[...]\nZhengyang Geng1,2,3, Yiyang Lu4,2,∗ Zongze Wu3 Eli Shechtman3 J. Zico Kolter1 Kaiming He2 1CMU 2MIT 3Adobe 4THU Equal contribution. Part of this work was done when Z. Geng was interning at Adobe and MIT, and when Y. Lu was interning at MIT.\n[...]\nMeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its “fastforward” nature introduces key challenges in both the training objective and the guidance mechanism. First, the original MF’s training target depends not only on the underlying ground-truth fields but also on the network itself. To address this issue, we recast the objective as a loss on the instantaneous velocity $v$ , re-parameterized by a network that predicts the average velocity $u$ . Our reformulation yields a more standard regression problem and improves the training stability. Second, the original MF fixes the classifier-free guidance scale during training, which sacrifices flexibility. We tackle this issue by formulating guidance as explicit conditioning variables, thereby retaining flexibility at test time. The diverse conditions are processed through in-context conditioning, which reduces model size and benefits performance. Overall, our improved MeanFlow (iMF) method, trained entirely from scratch, achieves 1.72 FID with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256. iMF substantially outperforms prior methods of t",
|
||
"source_url": "https://arxiv.org/abs/2512.02012",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2512.02012",
|
||
"_exa_published_date": "2025-12-01T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2511.19065] Understanding, Accelerating, and Improving MeanFlow Training",
|
||
"snippet": "Accelerating, and Improving MeanFlow Training\n[...]\n# Title:Understanding, Accelerating, and Improving MeanFlow Training\n\nAuthors: Jin-Young Kim, Hyojun Go, Lea Bogensperger, Julius Erbach, Nikolai Kalischek, Federico Tombari, Konrad Schindler, Dominik Narnhofer\n\nView PDF HTML (experimental)\n[...]\n> Abstract:MeanFlow promises high-quality generative modeling in few steps, by jointly learning instantaneous and average velocity fields. Yet, the underlying training dynamics remain unclear. We analyze the interaction between the two velocities and find: (i) well-established instantaneous velocity is a prerequisite for learning average velocity; (ii) learning of instantaneous velocity benefits from average velocity when the temporal gap is small, but degrades as the gap increases; and (iii) task-affinity analysis indicates that smooth learning of large-gap average velocities, essential for one-step generation, depends on the prior formation of accurate instantaneous and small-gap average velocities. Guided by these observations, we design an effective training scheme that accelerates the formation of instantaneous velocity, then shifts emphasis from short- to long-interval average velocity. Our enhanced MeanFlow training yields faster convergence and significantly better few-step generation: With the same DiT-XL backbone, our method reaches an impressive FID of 2.87 on 1-NFE ImageNet 256x256, compared to 3.43 for the conventional MeanFlow baseline. Alternatively, our method matche",
|
||
"source_url": "https://arxiv.org/abs/2511.19065",
|
||
"discovered_for": [
|
||
"rw.mean_flow"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2511.19065",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "RT-1: Robotics Transformer for Real-World Control at Scale",
|
||
"snippet": "By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project’s website and videos can be found at robotics-transformer1.github.io\n[...]\nThe second challenge lies in the design of the model itself. Effective robotic multi-task learning requires a high capacity model, and Transformer (Vaswani et al., 2017) models excel in this regard, particularly when it is necessary to learn many tasks conditioned, as in our case, on language instruct",
|
||
"source_url": "https://arxiv.org/html/2212.06817",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2212.06817",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "RT-1: Robotics Transformer for Real-World Control at Scale - arXiv",
|
||
"snippet": "By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project’s website and videos can be found at robotics-transformer1.github.io\n[...]\nThe second challenge lies in the design of the model itself. Effective robotic multi-task learning requires a high capacity model, and Transformer (Vaswani et al., 2017) models excel in this regard, particularly when it is necessary to learn many tasks conditioned, as in our case, on language instruct",
|
||
"source_url": "https://arxiv.org/abs/2212.06817",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2212.06817",
|
||
"_exa_published_date": "2022-12-13T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Bringing the RT-1-X Foundation Model to a SCARA robot",
|
||
"snippet": "Traditional robotic systems require specific training data for each task, environment, and robot form. While recent advancements in machine learning have enabled models to generalize across new tasks and environments, the challenge of adapting these models to entirely new settings remains largely unexplored. This study addresses this by investigating the generalization capabilities of the RT-1-X robotic foundation model to a type of robot unseen during its training: a SCARA robot from UMI-RTX.\n[...]\nRecent breakthroughs in machine learning and artificial intelligence suggest that training on large, diverse datasets can lead to highly\n[...]\nwhich often exceed\n[...]\nperformance of models developed for\n[...]\ndatasets tailored to\n[...]\nAs a result,\n[...]\nfield has been exploring more\n[...]\ncan adapt to\n[...]\n. Recent advancements like transformer\n[...]\n’s RT-1 [brohan_rt-1_2022], which demonstrate the potential for\n[...]\nGoogle’s RT-1 model is an impressive work, tested on a collection of real-world robotic experiences, where in different institutes a fleet of robots were performing 700 tasks [brohan_rt-1_2022]. The robots in the training set, such as the Franka, Kuka iiwa, UR5 and the EveryDay robot, can move their end-effector in a spherical working-space. None of the robots in the dataset is of the SCARA (Selective Compliance Assembly Robot Arm) type. With a SCARA robot the movement of z-axis is decoupled from the movement in the x-y plane, which gives a SCARA robot an kidney ",
|
||
"source_url": "https://arxiv.org/html/2409.03299v1",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2409.03299v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | OpenReview",
|
||
"snippet": "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | OpenReview\n[...]\n## RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nTL;DR: Vision-language models, trained on Internet-scale data, can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning.\n[...]\nAbstract: We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows th",
|
||
"source_url": "https://openreview.net/forum?id=XMQgwiJ7KSX",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=XMQgwiJ7KSX",
|
||
"_exa_published_date": "2023-08-30T15:38:01.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] RT-2: Vision-Language-Action Models Transfer Web Knowledge to ...",
|
||
"snippet": "Transfer Web Knowledge\n[...]\n# RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nWe study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot t",
|
||
"source_url": "https://arxiv.org/pdf/2307.15818",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2307.15818",
|
||
"_exa_published_date": "2023-07-28T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2307.15818] RT-2: Vision-Language-Action Models Transfer Web ...",
|
||
"snippet": "Transfer Web Knowledge\n[...]\n# RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nWe study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot t",
|
||
"source_url": "https://arxiv.org/abs/2307.15818",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2307.15818",
|
||
"_exa_published_date": "2023-07-28T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "OpenVLA: An Open-Source Vision-Language-Action Model - arXiv",
|
||
"snippet": "OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n# OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\nLarge policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challenges, we introduce OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations. OpenVLA builds on a Llama 2 language model combined with a visual encoder that fuses pretrained features from DINOv2 and SigLIP. As a product of the added data diversity and new model components, OpenVLA demonstrates strong results for generalist manipulation, outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters. We further show that we can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving multiple objects and strong language",
|
||
"source_url": "https://arxiv.org/abs/2406.09246",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2406.09246",
|
||
"_exa_published_date": "2024-06-13T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] OpenVLA: An Open-Source Vision-Language-Action Model | Semantic Scholar",
|
||
"snippet": "[PDF] OpenVLA: An Open-Source Vision-Language-Action Model | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Kim2024OpenVLAAO,\n title={OpenVLA: An Open-Source Vision-Language-Action Model},\n author={Moo Jin Kim and Karl Pertsch and Siddharth Karamcheti and Ted Xiao and Ashwin Balakrishna and Suraj Nair and Rafael Rafailov and Ethan Paul Foster and Grace Lam and Pannag R. Sanketi and Quan Vuong and Thomas Kollar and Benjamin Burchfiel and Russ Tedrake and Dorsa Sadigh and Sergey Levine and Percy Liang and Chelsea Finn},\n journal={ArXiv},\n year={2024},\n volume={abs/2406.09246},\n url={https://api.semanticscholar.org/CorpusID:270440391}\n}\n```",
|
||
"source_url": "https://www.semanticscholar.org/reader/8f9ceb5ffad8e7a066dfc9d9aaa5153b714740ee",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://www.semanticscholar.org/reader/8f9ceb5ffad8e7a066dfc9d9aaa5153b714740ee",
|
||
"_exa_published_date": "2024-06-13T09:02:12.000Z"
|
||
},
|
||
{
|
||
"title": "OpenVLA: An Open-Source Vision-Language-Action Model",
|
||
"snippet": "OpenVLA: An Open-Source Vision-Language-Action Model | OpenReview\n\n## OpenVLA: An Open-Source Vision-Language-Action Model\n\n### Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, Chelsea Finn \n\nCoRL 2024everyonesince 05 Sept 2024\">Everyone Revisions BibTeX CC BY 4.0\n\nKeywords: Vision-Language-Action Models, Generalist Policies, Large-scale Robot Learning, Robotic Manipulation, Robotics, Vision-Language Models\n\nTL;DR: We introduce OpenVLA, a state-of-the-art, open-source 7B-parameter VLA model that obtains strong performance for cross-embodiment robot control out-of-the-box and can be easily adapted to new robot setups via parameter-efficient fine-tuning.\n\nAbstract: Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challeng",
|
||
"source_url": "https://openreview.net/forum?id=ZMnD6QZAE6",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=ZMnD6QZAE6",
|
||
"_exa_published_date": "2024-09-05T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2406.09246v3] OpenVLA: An Open-Source Vision-Language-Action Model",
|
||
"snippet": "[2406.09246v3] OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n# Title:OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n> Abstract:Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challenges, we introduce OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations. OpenVLA builds on a Llama 2 language model combined with a visual encoder that fuses pretrained features from DINOv2 and SigLIP. As a product of the added data diversity and new model components, OpenVLA demonstrates strong results for generalist manipulation, outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters. We further show that we can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving mult",
|
||
"source_url": "https://arxiv.org/abs/2406.09246v3",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2406.09246v3",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Revision History for OpenVLA: An Open-Source... - OpenReview",
|
||
"snippet": "Revisions | OpenReview\n\nLoading",
|
||
"source_url": "https://openreview.net/revisions?id=ZMnD6QZAE6",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/revisions?id=ZMnD6QZAE6",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "A Vision-Language-Action Model for Affordable and Efficient Robotics",
|
||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||
"source_url": "https://arxiv.org/abs/2506.01844",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2506.01844",
|
||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable ... - arXiv",
|
||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||
"source_url": "https://arxiv.org/pdf/2506.01844",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2506.01844",
|
||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "SmolVLA: A vision-language-action model for affordable and ... - arXiv",
|
||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||
"source_url": "https://arxiv.org/html/2506.01844v1",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2506.01844v1",
|
||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics | Semantic Scholar",
|
||
"snippet": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Shukor2025SmolVLAAV,\n title={SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics},\n author={Mustafa Shukor and Dana Aubakirova and Francesco Capuano and Pepijn Kooijmans and Steven Palma and Adil Zouitine and Michel Aractingi and Caroline Pascal and Martino Russi and Andr{\\'e}s Marafioti and Simon Alibert and Matthieu Cord and Thomas Wolf and R{\\'e}mi Cad{\\`e}ne},\n journal={ArXiv},\n year={2025},\n volume={abs/2506.01844},\n url={https://api.semanticscholar.org/CorpusID:279119427}\n}\n```",
|
||
"source_url": "https://www.semanticscholar.org/reader/6ab4d113676d00e74b55e918fee4c7affaa8652f",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://www.semanticscholar.org/reader/6ab4d113676d00e74b55e918fee4c7affaa8652f",
|
||
"_exa_published_date": "2025-06-02T13:54:14.000Z"
|
||
},
|
||
{
|
||
"title": "Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots",
|
||
"snippet": "By leveraging NF4 quantization and the llama-cpp runtime, the proposed LiteVLA implementation pioneers the CPU-only deployment path, achieving functional asynchronous visuomotor control on the low-cost Raspberry Pi 4. This represents a novel deployment strategy not demonstrated by prior GPU-centric VLA frameworks such as SmolVLA by Shukor et al. [22], whose work focused primarily on static robotic arms. Beyond proving technical feasibility, this work establishes a scalable methodology for deploying generalist robot intelligence under strict computational budgets.\n[...]\nParameter-efficient adaptation. We fine-tune a compact SmolVLM backbone using LoRA (rank 8, $\\alpha{=}8$ , dropout 0.1) to specialize visuomotor translation under tight memory/compute budgets (Alg. 1; Sec. III, pp. 2–3).\n[...]\n. 4).\n[...]\nLarge-scale multimodal systems such as PaLM-E, SayCan, and RT-2 have shown that unified language-conditioned reasoning enables robots to follow natural language commands and execute complex manipulation tasks. However, these approaches rely heavily on cloud-based computation and high-end GPUs, making them impractical for resource-limited or field-deployed robots. SMolVLA by Shukor et al. [22] introduced a small and efficient vision-language-action framework designed for community-driven robotic experimentation. It demonstrated that compact multimodal transformers could achieve competitive visuomotor reasoning performance while running on consumer-grade GPUs or CPUs. Nonetheles",
|
||
"source_url": "https://arxiv.org/html/2511.05642",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2511.05642",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "A Vision-Language-Action Flow Model for General Robot Control",
|
||
"snippet": "𝜋₀: A Vision-Language-Action Flow Model for General Robot Control\n[...]\n# $\\pi_{0}$ : A Vision-Language-Action Flow Model for General Robot Control\n[...]\nholds tremendous promise to unlock\n[...]\nfull potential of flexible, general, and dexterous\n[...]\nsystems, as well as to address some of\n[...]\ndeepest questions in artificial intelligence\n[...]\nHowever, bringing robot learning to\n[...]\nlevel of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how\n[...]\ncan design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of\n[...]\nability to perform tasks via direct prompting, follow language instructions from people and from a high-level VLM policy, and\n[...]\nability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.\n[...]\nIn this paper, we present a prototype model and learning framework, which we call $\\pi_",
|
||
"source_url": "https://arxiv.org/abs/2410.24164",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2410.24164",
|
||
"_exa_published_date": "2024-10-31T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "π: A Vision-Language-Action Flow Model for General Robot Control | OpenReview",
|
||
"snippet": "π: A Vision-Language-Action Flow Model for General Robot Control | OpenReview\n[...]\n## π: A Vision-Language-Action Flow Model for General Robot Control\n[...]\nAbstract: Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.",
|
||
"source_url": "https://openreview.net/forum?id=38a45ho9Nq",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=38a45ho9Nq",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Efficient Action Tokenization for Vision-Language-Action Models",
|
||
"snippet": "Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, such models require us to choose a tokenization of our continuous action signals, which determines how the discrete symbols predicted by the model map to continuous robot actions. We find that current approaches for robot action tokenization, based on simple per-dimension, per-timestep binning schemes, typically perform poorly when learning dexterous skills from high-frequency robot data. To address this challenge, we propose a new compression-based tokenization scheme for robot actions, based on the discrete cosine transform. Our tokenization approach, Frequency-space Action Sequence Tokenization (FAST), enables us to train autoregressive VLAs for highly dexterous and high-frequency tasks where standard discretization methods fail completely. Based on FAST, we release FAST+, a universal robot action tokenizer, trained on 1M real robot action trajectories. It can be used as a black-box tokenizer for a wide range of robot action sequences, with diverse action spaces and control frequencies. Finally, we show that, when combined with the $\\bm{\\pi_{0}}$ VLA, our method can scale to training on 10k hours of robot data and match the performance of diffusion VLAs, while reducing training time by up to 5x.\n[...]\nFigure 4: Overview of the FAST action tokenization pipeline. Given a normalized c",
|
||
"source_url": "https://arxiv.org/abs/2501.09747",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2501.09747",
|
||
"_exa_published_date": "2025-01-16T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "FAST: Efficient Action Tokenization for Vision-Language ... - arXiv",
|
||
"snippet": "Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, such models require us to choose a tokenization of our continuous action signals, which determines how the discrete symbols predicted by the model map to continuous robot actions. We find that current approaches for robot action tokenization, based on simple per-dimension, per-timestep binning schemes, typically perform poorly when learning dexterous skills from high-frequency robot data. To address this challenge, we propose a new compression-based tokenization scheme for robot actions, based on the discrete cosine transform. Our tokenization approach, Frequency-space Action Sequence Tokenization (FAST), enables us to train autoregressive VLAs for highly dexterous and high-frequency tasks where standard discretization methods fail completely. Based on FAST, we release FAST+, a universal robot action tokenizer, trained on 1M real robot action trajectories. It can be used as a black-box tokenizer for a wide range of robot action sequences, with diverse action spaces and control frequencies. Finally, we show that, when combined with the $\\bm{\\pi_{0}}$ VLA, our method can scale to training on 10k hours of robot data and match the performance of diffusion VLAs, while reducing training time by up to 5x.\n[...]\nFigure 4: Overview of the FAST action tokenization pipeline. Given a normalized c",
|
||
"source_url": "https://arxiv.org/html/2501.09747v1",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2501.09747v1",
|
||
"_exa_published_date": "2025-01-16T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "ActionCodec: What Makes for Good Action Tokenizers",
|
||
"snippet": "without any robotics\n[...]\nintroduce ActionCodec, a robust action tokenizer that integrates the\n[...]\n. Moreover, ActionCodec leverages Residual Vector Quantization (RVQ) (Lee et al., 2022) post-training to refine reconstruction fidelity and incorporates embodiment-specific soft prompts to facilitate knowledge transfer across diverse robotic platforms\n[...]\nthat VLA\n[...]\nActionCodec, without any additional architectural modifications\n[...]\nefficiency, success rates\n[...]\n. ActionCodec achieves SOTA performance in both\n[...]\nenvironments, providing a systematic\n[...]\nfor the future of VQ-\n[...]\nOur contributions are\n[...]\nas follows:\n[...]\nTokenization Schemes\n[...]\n20\n[...]\nsuffers from low training efficiency and ignores the\n[...]\nparallel decoding (\n[...]\n(Goy\n[...]\n., 2025),\n[...]\nfundamental inefficiencies of heuristic binning. Other\n[...]\nrepresent actions as strings for direct VLM prediction (Hancock et\n[...]\nhowever, this approach\n[...]\nno significant performance benefits while greatly increasing the token budget and extending latency to several seconds, limiting\n[...]\nPertsch et al., 2025) introduces Byte-Pair Encoding (BPE) on frequency-domain signals; however, its reliance on fixed geometric priors limits its capacity for cross-embodiment knowledge transfer. Data-driven schemes, particularly those based on Vector Quantization (VQ) (Wang et al., 2025b; Belkhale and Sadigh, 2024; Mete et al., 2024; Lee et al., 2024), offer a more flexible alternative by learning disc",
|
||
"source_url": "https://arxiv.org/abs/2602.15397",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2602.15397",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "OAT: Ordered Action Tokenization",
|
||
"snippet": "action tokenization\n[...]\nTo bridge this gap, we introduce Ordered Action Tokenization (OAT), a learned action tokenizer that discretizes continuous action chunks into highly compressed and causally ordered token sequences. OAT employs transformer-based register tokens to aggregate temporal information, finite scalar quantization (FSQ) to construct a discrete bottleneck, and nested dropout to explicitly induce ordering that aligns the latent space with autoregressive generation. The resulting tokenization ensures that any token prefix corresponds to a plausible action chunk. Beyond improved modelability, the ordered structure learned by OAT enables a key capability absent from prior approaches: prefix-based decoding. Autoregressive policies may terminate generation early and still produce valid actions, yielding a natural trade-off between computation and action fidelity. As additional tokens are generated, decoded actions are progressively refined.\n[...]\nAn alternative line of work explores frequency-domain compression, for instance Frequency-space Action Sequence Tokenization (FAST) [49], which employs the Discrete Cosine Transform (DCT) to decompose action chunks into frequency coefficients, followed by Byte Pair Encoding (BPE) [18]. FAST achieves high information density (P.1), and crucially, its low-frequency components first then high-frequency components ordering (P.3) improves downstream autoregressive policies: early token predictions capture the overall trajectory s",
|
||
"source_url": "https://arxiv.org/html/2602.04215",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2602.04215",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding",
|
||
"snippet": "the above challenges, we present a novel parallel decoding framework for the mainstream VLA model with action chunking, called Parallel\n[...]\nfor VLA (PD-VLA). Fig. 1 illustrates the\n[...]\nconcept of our parallel decoding approach. Our\n[...]\naction decoding as a system of\n[...]\nsolved through parallel fixed-point iteration methods, e.g., Jacobi fix-point iteration method [32]. This approach preserves\n[...]\nimproving decoding speed\n[...]\nthat we only accelerate the decoding process\n[...]\nVLA inference.\n[...]\n, our method enables friendly\n[...]\ntraining-free acceleration without redesign and modification of models (\n[...]\n, our method\n[...]\nsynergy with existing acceleration\n[...]\nVarious acceleration strategies, including quantization [21] and token pruning [5], have been effectively applied to LLMs, yet they often fail to meet the stringent real-time requirements of action generation. Efforts to enhance efficiency have led to architectural modifications in VLA models, such as DeeR-VLA [43], which dynamically adjusts inference depth, and QAIL [33], which integrates quantization-aware training. Further innovations, like RoboMamba [25] and TinyVLA [41], replace traditional attention mechanisms or focus on developing lightweight models from the ground up, frequently necessitating model re-training and additional data collection. Meanwhile, VLA-Cache [42] selectively caches static tokens and recomputes only dynamic or task-relevant ones. FAST [34] proposes a compression-based toke",
|
||
"source_url": "https://arxiv.org/html/2503.02310v2",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2503.02310v2",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware | OpenReview",
|
||
"snippet": "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware | OpenReview\n[...]\n## Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware\n[...]\nAbstract: Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations.",
|
||
"source_url": "https://openreview.net/forum?id=e8Eu1lqLaf",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=e8Eu1lqLaf",
|
||
"_exa_published_date": "2023-07-09T06:52:40.000Z"
|
||
},
|
||
{
|
||
"title": "Learning Fine-Grained Bimanual Manipulation with Low-Cost ... - arXiv",
|
||
"snippet": "Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations. Project website: tonyzhaozh.github.io/aloha\n[...]\nImitation learning algorithm. Tasks that require precision and visual feedback present a significant challenge for imitation learning, even with high-quality demonstrations. Small errors in the predicted action can incur large differences in the state, exacerbating the “co",
|
||
"source_url": "https://arxiv.org/abs/2304.13705",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2304.13705",
|
||
"_exa_published_date": "2023-04-23T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with Transformers",
|
||
"snippet": "coordinated biman\n[...]\n. To address the\n[...]\narms, particularly for synchronized actions. Therefore, we propose a novel imitation learning architecture that predicts cooperative actions. We differentiate the architecture for both arms and add an intermediate encoder layer, Inter-Arm Coordinated transformer Encoder (IACE),\n[...]\nInter-Arm Coordinated transformer Encoder (IACE), that can adjust the synchronization and timing of potential bimanual movements against the encoders corresponding to each arm. Our overall model\n[...]\na local Transformer encoder for each\n[...]\narm trajectory, the IACE to facilitate learning biman\n[...]\nactions, and a Transformer decoder to\n[...]\nthe action chunk. We compare two types of Transformer decoders: split decoders and single decoders.\n[...]\nWe build our proposed models on the ACT model to design different encoder and decoder structures. In particular, we propose a new design called the inter-arm coordinated transformer Encoder (IACE), which helps synchronize and time the movements of both arms.\n[...]\npropose basic architectures that consist of encoders\n[...]\narm, designed to leverage the biman\n[...]\nfeatures the IACE, allowing the individual robot arms to learn their trajectories while simultaneously considering the state\n[...]\nThe model should focus on the corresponding wrist camera and joint values to determine the appropriate trajectory for each robot arm. Each arm is supported by its local encoder. Global information is also integrated t",
|
||
"source_url": "https://arxiv.org/html/2503.13916",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2503.13916",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "ALPHA-𝛼 and Bi-ACT Are All You Need: Importance of Position and Force Information/Control for Imitation Learning of Unimanual and Bimanual Robotic Manipulation with Low-Cost System",
|
||
"snippet": "Autonomous manipulation in everyday tasks requires flexible action generation to handle complex, diverse real-world environments, such as objects with varying hardness and softness. Imitation Learning (IL) enables robots to learn complex tasks from expert demonstrations. However, a lot of existing methods rely on position/unilateral control, leaving challenges in tasks that require force information/control, like carefully grasping fragile or varying-hardness objects. As the need for diverse controls increases, there are demand for low-cost bimanual robots that consider various motor inputs. To address these challenges, we introduce Bilateral Control-Based Imitation Learning via Action Chunking with Transformers(Bi-ACT) and”A” ”L”ow-cost ”P”hysical ”Ha”rdware Considering Diverse Motor Control Modes for Research in Everyday Bimanual Robotic Manipulation (ALPHA- $\\alpha$ ). Bi-ACT leverages bilateral control to utilize both position and force information, enhancing the robot’s adaptability to object characteristics such as hardness, shape, and weight. The concept of ALPHA- $\\alpha$ is affordability, ease of use, repairability, ease of assembly, and diverse control modes (position, velocity, torque), allowing researchers/developers to freely build control systems using ALPHA- $\\alpha$ . In our experiments, we conducted a detailed analysis of Bi-ACT in unimanual manipulation tasks, confirming its superior performance and adaptability compared to Bi-ACT without force control. Base",
|
||
"source_url": "https://arxiv.org/html/2411.09942",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2411.09942",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Learning Fine-Grained Bimanual Manipulation with Low-Cost ...",
|
||
"snippet": "�W7/��A�/�\u0003c�s^���u�\"N�f�|5\n[...]\nǨ,��~.��a(Ц�\u0007\u001e�Pb��%��\u0019��\u001a?�3���\u001eS��~����l���\u0012�j\u0012tf;m���3\n[...]\n��7��h�\n[...]\n4\u001e\u0013L\n[...]\n\u000f9\\����� \u0001\u001a�� ��'��T�G�15�f����-ZB3���4������\u0000�[��\u0006����]�ӦbQi�uK\u0003��g4� h⠯� �8�@S�B��\u0002˔.�X��\u0007��IJ)��H\u0007I�����A��\u0011W�d������\u001a��\n[...]\n�Z\u0002O�[�\u0018^j�M\u000f\u001d��wE��df�\u0014h��C\u0004�\\���xͫ\u0010SB�\b��7�\u001aNb�ܫ#~8�7MD\n[...]\n�\u0001�\n[...]\nʬ`��y�1�1\u001fBf���R\u001fT$,Υ;lv\u0015@+%�\u0000�O�;k�S�6\"�\u0002#�J\n[...]\n�B2�+�\u0017.�\u0004�i\u0015�C �\u0005 ��\u0003�a�G�T�\u001b�W�\u0015\u000eג��RM�;�\u0004?�9$\u0014&�\u0013zT��Ʌ��z=7-�ip��1S3�� ���+\u001fI\u0013��q^��\u0018:���.�3,�\"t�v�\u001c��(c2�M�f�~�\u0010'\u001c�ӝ�ӨBf�d&\u0004�\u001d���\u00042�\b� C\u001e\u0006SΩ�T�A�8M!Ulŏ�f-�\u0003\u0004ݩF�30�nR�6ϛ�]E\b�\u0003]\\r�\\�n�5�� Q�uϬZZ|\u0003S�D<�_(�DU�<�8�0�x�)\u0006p��h�T\u0016�L\u0006��ו�X��#�n^+^q�8e0\n[...]\n�U*c��U�OB\u0011:\u0013 *��. �D:\u001b��iځ\u0015 l:��&:�Y��bxހl̰���?���\u00149�M �^|�ʛ1H���6����|��p��37I\u0002\u0001Z棍�\u0010�B���!���\\�A\u0012K��d�r�T5�>�ƨ^�t\u001e���\\|�m\u0011�>poŅ���\u0012��#5�%w[�|+\u0016iX���fy���X�ޕ�ݔs��M�\u0014��:�9���1b�����7�\u0015=U\n[...]\n\u0012�m\u001d�����!� I4\n[...]\n��A<����䬆��ᯟ \u0006�\u0018&�x⥌�;���H� �,#n�\u0013��U?���x�ؔ��d�\u0001��]�8�HoJ�}�\u0007 �ܓ �=��Z�\u0005G��T?�\u0004㘨�\u0007\u001a�!��*ǻ�5��[��x�D��i\u001c�Gc��\\\u000f�%��2��c�\u001b�d��6=d��Sm�;��\u0004��ݏF\u001e�\b1T�\u0006 �\u0012ch\u0007��� \u000f�\bV^��ԥ�e��_P�,� \u0010h_�O����\u0012ٱj\u0004��J\u001f?\u0012\u0012y\u0017 :{Zs�\u001b�x\u001d \\_�W�!\u0006���Ь��p���U!���\u001d�lm�n�B��E�\u0018g��\u0012\u0014w�ɺ��8�\u0010��N�\u0011v�M����`�\u001dYD|��=ZI8E��\u001c~\\�Ϫ��1���K���(\u0006�\u0005�����&\u0015\u0010R�\u0012_\u001e��\u0000��F�{ ��\u001bm�&��u\\�\u0005��ݶ�Ng��\u0011��4!��\u0005�Kh��.��a�3o�.'���2\\��wi3� >)iJ��sR� 3 �FB��O��\u0018�7��\u0019劍o s,���*��P\u0012\u001e\u0018\"���`�G�\u000f�2��l$q��Y��O?r\u001e_�P{\u0001�ܷG\b ����`�6\u001f��z��.��+���m��t%9��#E!$JtJe�$8C6�ͥ�\u001e�Ù�OV�]��bnvi)u0�fHy'���\u0010a����_,�D�M&uǒ�|�s]�D%�E>\"\"6� ��$ �� A����C��Q\u0003�`֩������}��\u0011� ��K/b�D��:�З�G<�\u001cy+wj�|U�\u0000r�\u0016Le�5��G��N|\u0010�K���� �Yv`\u0000K� �J0",
|
||
"source_url": "https://openreview.net/pdf/4abc35d9793e56c5b73634eaf903e2495311fbcf.pdf",
|
||
"discovered_for": [
|
||
"rw.vla"
|
||
],
|
||
"_exa_id": "https://openreview.net/pdf/4abc35d9793e56c5b73634eaf903e2495311fbcf.pdf",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[1512.03385] Deep Residual Learning for Image Recognition - arXiv",
|
||
"snippet": "[1512.03385] Deep Residual Learning for Image Recognition\n[...]\n# Deep Residual Learning for Image Recognition\n[...]\nKaiming He Xiangyu Zhang Shaoqing Ren Jian Sun Microsoft Research {kahe, v-xiangz, v-shren, jiansun}@microsoft.com\n[...]\nDeeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers—8 $\\times$ deeper than VGG nets [41] but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers.\n[...]\nreferenced mapping. To\n[...]\nextreme, if\n[...]\nwould be easier\n[...]\nresidual to zero than\n[...]\nnonlinear layers.\n[...]\non ImageNet [36\n[...]\nOn the ImageNet classification dataset [36], we obtain excellent results by extremely deep residual nets. Our 152-layer residual net is the deepest network ever presented on ImageNet, while still having lower complexity than VGG nets [41]. Our ensem",
|
||
"source_url": "https://arxiv.org/abs/1512.03385",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1512.03385",
|
||
"_exa_published_date": "2015-12-10T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[1512.03385] Deep Residual Learning for Image Recognition",
|
||
"snippet": "[1512.03385] Deep Residual Learning for Image Recognition\n[...]\n# Deep Residual Learning for Image Recognition\n[...]\nKaiming He Xiangyu Zhang Shaoqing Ren Jian Sun Microsoft Research {kahe, v-xiangz, v-shren, jiansun}@microsoft.com\n[...]\nDeeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers—8 $\\times$ deeper than VGG nets [41] but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers.\n[...]\nreferenced mapping. To\n[...]\nextreme, if\n[...]\nwould be easier\n[...]\nresidual to zero than\n[...]\nnonlinear layers.\n[...]\non ImageNet [36\n[...]\nOn the ImageNet classification dataset [36], we obtain excellent results by extremely deep residual nets. Our 152-layer residual net is the deepest network ever presented on ImageNet, while still having lower complexity than VGG nets [41]. Our ensem",
|
||
"source_url": "https://arxiv.org/abs/1512.03385v1",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1512.03385v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Deep Residual Learning for Image Recognition - People | MIT CSAIL",
|
||
"snippet": "Deep Residual Learning\nfor Image Recognition\nKaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun\nwork done at\nMicrosoft Research Asia\n[...]\nResNet @ ILSVRC & COCO 2015 Competitions\n[...]\n1st places in all five main tracks\n[...]\n• ImageNet Classification: “Ultra-deep” 152-layer nets \n• ImageNet Detection: 16% better than 2nd\n[...]\n• ImageNet Localization: 27% better than 2nd\n[...]\n• COCO Detection: 11% better than 2nd\n[...]\n• COCO Segmentation: 12% better than 2nd\n[...]\n*improvements are relative numbers\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nRevolution of Depth\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiang",
|
||
"source_url": "https://pdfs.semanticscholar.org/1cea/9b1931b9e87641708fec43d03f2a58f4d2b0.pdf",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://pdfs.semanticscholar.org/1cea/9b1931b9e87641708fec43d03f2a58f4d2b0.pdf",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Deep Residual Learning for Image Recognition: A Survey - MDPI",
|
||
"snippet": "Deep Residual Learning for Image Recognition: A Survey\n[...]\n# Deep Residual Learning for Image Recognition: A Survey\n[...]\nMuhammad Shafiq\n[...]\n1,* and\n\nZhaoquan Gu\n[...]\n2,3,*\n[...]\nCyberspace Institute of Advanced Technology, Guangzhou University, Guangzhou 510006, China\n[...]\nDepartment of New Networks, Peng Cheng Laboratory, Shenzhen 518055, China\n[...]\nDepartment of Computer Science and Technology, Harbin Institute of Technology, Shenzhen 518055, China\n[...]\nAppl. Sci. 2022, 12(18), 8972; https://doi.org/10.3390/app12188972\n[...]\nDeep Residual Networks have recently been shown to significantly improve the performance of neural networks trained on ImageNet, with results beating all previous methods on this dataset by large margins in the image classification task. However, the meaning of these impressive numbers and their implications for future research are not fully understood yet. In this survey, we will try to explain what Deep Residual Networks are, how they achieve their excellent results, and why their successful implementation in practice represents a significant advance over existing techniques. We also discuss some open questions related to residual learning as well as possible applications of Deep Residual Networks beyond ImageNet. Finally, we discuss some issues that still need to be resolved before deep residual learning can be applied on more complex problems.\n[...]\ndeep residual learning for image recognition\n[...]\ndeep residual learning;\n[...]\nDeep resid",
|
||
"source_url": "https://www.mdpi.com/2076-3417/12/18/8972",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://www.mdpi.com/2076-3417/12/18/8972",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[1603.05027] Identity Mappings in Deep Residual Networks - arXiv",
|
||
"snippet": "Kaiming He Xiangyu Zhang Shaoqing Ren Jian Sun\n[...]\nDeep residual networks [1] have emerged as a family of extremely deep architectures showing compelling accuracy and nice convergence behaviors. In this paper, we analyze the propagation formulations behind the residual building blocks, which suggest that the forward and backward signals can be directly propagated from one block to any other block, when using identity mappings as the skip connections and after-addition activation. A series of ablation experiments support the importance of these identity mappings. This motivates us to propose a new residual unit, which makes training easier and improves generalization. We report improved results using a 1001-layer ResNet on CIFAR-10 (4.62% error) and CIFAR-100, and a 200-layer ResNet on Image\n[...]\n. Code is available at: https://github.com/KaimingHe/resnet-1k-layers.\n[...]\n- [1] He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. (2016)",
|
||
"source_url": "https://arxiv.org/abs/1603.05027",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1603.05027",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[2603.15031] Attention Residuals - arXiv",
|
||
"snippet": "Untitled Document\n\n$0$ $5$ $10$ $15$ $20$\n\nUntitled Document\n$0$ $5$ $10$ $15$ $20$\nBETA",
|
||
"source_url": "https://arxiv.org/abs/2603.15031",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2603.15031",
|
||
"_exa_published_date": "2026-03-16T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm",
|
||
"snippet": "In this paper, we propose SiameseNorm, an elegant two\n[...]\nstream residual architecture that unifies\n[...]\nmaintain two residual streams\n[...]\nshared parameters:\n[...]\nthe advantages of\n[...]\nnegligible computational overhead\n[...]\nboosts accuracy from\n[...]\n128\n[...]\n639.6\n[...]\nwhere the product denotes an ordered composition of Jacobians from layer N−1N-1 down to i+1i+1. Notably, the term 𝐈\\mathbf{I} corresponds to the skip connection, which preserves an explicit identity gradient path. This allows gradients to flow through the network without explicit attenuation, facilitating the training of large scale models. However, it implicitly allows the representation magnitudes to grow unbounded. As noted previously, Pre-Norm exhibits insufficient effective depth, an issue that likely stems from a structural mismatch: As shown in Figure˜2(a), the main path accumulates residual updates without re-normalization, causing hidden state magnitudes to grow with depth (peri-ln). Consequently, deeper blocks encounter a scaling imbalance: they must influence an increasingly high-magnitude main path while being restricted to normalized, fixed-scale inputs. This growing disparity effectively dilutes the relative contribution of deeper layers, thereby limiting the effective depth of the model.\n[...]\nBy maintaining a clean identity path, Pre-Norm ensures stable gradient propagation. However, this comes at the cost of unbounded magnitude growth. As illustrated in Figure 2(a), while the input ",
|
||
"source_url": "https://arxiv.org/html/2602.08064v1",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/html/2602.08064v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] Attention Residuals - arXiv",
|
||
"snippet": "Untitled Document\n\n$0$ $5$ $10$ $15$ $20$\n\nUntitled Document\n$0$ $5$ $10$ $15$ $20$\nBETA",
|
||
"source_url": "https://arxiv.org/pdf/2603.15031",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/2603.15031",
|
||
"_exa_published_date": "2026-03-16T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[2409.19606] Hyper-Connections - arXiv",
|
||
"snippet": "We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed in residual connection variants, such as the seesaw effect between gradient vanishing and representation collapse. Theoretically, hyper-connections allow the network to adjust the strength of connections between features at different depths and dynamically rearrange layers. We conduct experiments focusing on the pre-training of large language models, including dense and sparse models, where hyper-connections show significant performance improvements over residual connections. Additional experiments conducted on vision tasks also demonstrate similar improvements. We anticipate that this method will be broadly applicable and beneficial across a wide range of AI problems.\n[...]\nDeep learning has achieved tremendous success across various domains, where residual connections (He et al., 2016) have been instrumental in contemporary neural network architectures, including transformers and CNNs. Residual connections help mitigate the problem of gradient vanishing, enabling the effective training of very deep networks. However, it is important to acknowledge that residual connections are not infallible solutions and still present limitations that remain unresolved.\n[...]\nDriven by the limitations of residual connections, an important question arises: Can neural networks autonomously learn the optimal streng",
|
||
"source_url": "https://arxiv.org/abs/2409.19606",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/2409.19606",
|
||
"_exa_published_date": "2024-09-29T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Hyper-Connections - OpenReview",
|
||
"snippet": "Hyper-Connections | OpenReview\n\n## Hyper-Connections\n\n### Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou\n\nICLR 2025 Postereveryonesince 04 Oct 2024\">Everyone Revisions BibTeX CC BY 4.0\n\nKeywords: Network Architecture, Residual Connections, LLMs, Pre-training\n\nAbstract: We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed in residual connection variants, such as the seesaw effect between gradient vanishing and representation collapse. Theoretically, hyper-connections allow the network to adjust the strength of connections between features at different depths and dynamically rearrange layers. We conduct experiments focusing on the pre-training of large language models, including dense and sparse models, where hyper-connections show significant performance improvements over residual connections. Additional experiments conducted on vision tasks also demonstrate similar improvements. We anticipate that this method will be broadly applicable and beneficial across a wide range of AI problems.\n\nPrimary Area: foundation or frontier models, including LLMs\n\nCode Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics.\n\nSubmission Guidelines: I certify that this submission complies with the submission instructions as described on https://iclr.cc/Con",
|
||
"source_url": "https://openreview.net/forum?id=9FqARW7dwB",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=9FqARW7dwB",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "Birkhoff-Exact Hyper-Connections: Exact Spectral Stability for Deep Residual Networks | OpenReview",
|
||
"snippet": "## Birkhoff-Exact Hyper-Connections: Exact Spectral Stability for Deep Residual Networks\n[...]\nKeywords: doubly stochastic matrices, spectral stability, deep residual networks, Birkhoff-von Neumann theorem, hyper-connections, token mixing, extreme depth training, quantization robustness\n[...]\nTL;DR: We propose BE-HC, which uses the Birkhoff-von Neumann theorem to construct exactly doubly stochastic mixing matrices as convex combinations of permutation matrices, enabling stable training at 1000+ layers where prior methods fail.\n[...]\nAbstract: Learnable information routing in deep networks faces the *depth-stability-efficiency trilemma*: architectures that scale to extreme depths often sacrifice efficiency; efficient approaches lack stability guarantees. Prior work uses iterative Sinkhorn-Knopp normalization to approximate doubly stochastic mixing matrices, but residual errors destabilize training beyond several hundred layers. We propose **Birkhoff-Exact Hyper-Connections (BE-HC)**, which leverages the Birkhoff-von Neumann theorem to construct *exactly* doubly stochastic matrices as convex combinations of permutation matrices. This guarantees spectral radius $\\rho = 1$ exactly—not approximately—enabling stable training at unprecedented depths. **Key results:** (1) *Extreme depth:* BE-HC trains stably at **1000 layers**, achieving 35.71% accuracy where ReZero and other baselines fail to converge. (2) *Long context:* BE-HC handles **8K tokens** on a single V100 GPU (22.56% vali",
|
||
"source_url": "https://openreview.net/forum?id=jpIjkN1B1Q",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=jpIjkN1B1Q",
|
||
"_exa_published_date": "2026-03-02T22:01:32.000Z"
|
||
},
|
||
{
|
||
"title": "Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections",
|
||
"snippet": "Multi-stream transformer architectures have recently been proposed as a promising direction for managing representation collapse and the vanishing gradient problem for residual connections, yet their internal mechanisms remain unexplored. In particular, the recently introduced Manifold-Constrained Hyper-Connections (mHC) architecture posits multiple residual streams with constrained interaction, but lacks in-depth mechanistic analysis. We present the first open-source mHC language model (https://huggingface.co/wgpeng/mhc-780m) and analyze the multiple-stream architecture with a suite of representation-level metrics and causal interventions to probe how parallel streams encode and utilize information. Specifically, we introduce a systematic stream ablation-and-rescue framework that enables direct causal comparison of residual streams during inference. Through targeted pairwise interventions and controlled recovery experiments, we distinguish functional redundancy from asymmetric utilization and reveal how information is distributed across streams beyond what is observable from representational similarity alone.\n[...]\nHyper-Connections extend the standard transformer residual architecture by allowing multiple residual streams per layer, dynamically mixed through learned routing matrices (He et al., 2015; Zhu et al., 2025). Manifold-Constrained Hyper-Connections (mHC) further refines this framework by imposing geometric constraints on inter-stream mixing (Xie et al., 2026).\n[...",
|
||
"source_url": "https://www.arxiv.org/pdf/2603.14833",
|
||
"discovered_for": [
|
||
"rw.attnres"
|
||
],
|
||
"_exa_id": "https://www.arxiv.org/pdf/2603.14833",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[1711.05101] Decoupled Weight Decay Regularization - arXiv",
|
||
"snippet": "[1711.05\n[...]\npled Weight Decay Regularization\n[...]\nIlya Loshchilov & Frank Hutter University of Freiburg Freiburg, Germany, {ilya,fh}@cs.uni-freiburg.de\n[...]\nL2 regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is not the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L2 regularization (often calling it “weight decay” in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by decoupling the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam’s generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at https://github.com/loshchil/AdamW-and-SGDW\n[...]\nThe main contribution of this paper is to improve regularization in Adam by decoupling ",
|
||
"source_url": "https://arxiv.org/abs/1711.05101",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1711.05101",
|
||
"_exa_published_date": "2017-11-14T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "[PDF] Decoupled Weight Decay Regularization - arXiv",
|
||
"snippet": "DECOUPLED WEIGHT DECAY REGULARIZATION\n[...]\nIlya Loshchilov & Frank Hutter\n[...]\nL2 regularization and weight decay regularization are equivalent for standard\n[...]\nstochastic gradient descent (when rescaled by the learning rate), but as we demon\u0002strate this is not the case for adaptive gradient algorithms, such as Adam. While\n[...]\nexpose), we propose a simple modification to recover the original formulation of\n[...]\nweight decay regularization by decoupling the weight decay from the optimization\n[...]\nsteps taken w.r.t. the loss function. We provide empirical evidence that our pro\u0002posed modification (i) decouples the optimal choice of weight decay factor from\n[...]\nthe setting of the learning rate for both standard SGD and Adam and (ii) substan\u0002tially improves Adam’s generalization performance, allowing it to compete with\n[...]\nSGD with momentum on image classification datasets (on which it was previously\n[...]\ntypically outperformed by the latter). Our proposed decoupled weight decay has\n[...]\ncommunity has implemented\n[...]\nit in TensorFlow and PyTorch; the complete source code for our experiments is\n[...]\navailable at https://github.com/loshchil/AdamW-and-SGDW\n[...]\nThe main contribution of this paper is to improve regularization in Adam by decoupling the weight\n[...]\ndecay from the gradient-based update. In a comprehensive analysis, we show that Adam generalizes\n[...]\nsubstantially better with decoupled weight decay than with L2 regularization, achieving 15% relative\n[.",
|
||
"source_url": "https://arxiv.org/pdf/1711.05101",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://arxiv.org/pdf/1711.05101",
|
||
"_exa_published_date": "2019-01-04T00:00:00.000Z"
|
||
},
|
||
{
|
||
"title": "Decoupled Weight Decay Regularization - OpenReview",
|
||
"snippet": "Decoupled Weight Decay Regularization | OpenReview\n\n## Decoupled Weight Decay Regularization\n\nICLR 2019 Conference Blind SubmissionReaders: Everyone\n\nAbstract: L$_2$ regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \\emph{not} the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L$_2$ regularization (often calling it ``weight decay'' in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by \\emph{decoupling} the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam's generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at \\url{https://github.com/loshchil/AdamW-and-SGDW}\n\nKeywords: optimization, regularization, weight decay, Adam\n\nCode: ",
|
||
"source_url": "https://openreview.net/forum?id=Bkg6RiCqY7",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://openreview.net/forum?id=Bkg6RiCqY7",
|
||
"_exa_published_date": "2018-09-27T08:45:35.000Z"
|
||
},
|
||
{
|
||
"title": "PyTorch: An Imperative Style, High-Performance Deep Learning Library",
|
||
"snippet": "PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\nDeep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several commonly used benchmarks.",
|
||
"source_url": "https://papers.nips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://papers.nips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[PDF] An Imperative Style, High-Performance Deep Learning Library - NIPS",
|
||
"snippet": "PyTorch: An Imperative Style, High-Performance Deep Learning Library\n\n| | Adam | Paszke | | Sam | Gross | | | Francisco | Massa | | |\n[...]\nAbstract Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several common benchmarks.\n[...]\nWith the increased interest in deep learning in recent years, there has been an explosion of machine learning tools. Many popular frameworks such as Caffe [1], CNTK [2], TensorFlow [3], and Theano [4], construct a static dataflow graph that represents the computation and which can then be applied repeatedly to batches of data. This approach provides visibility into the whole comput",
|
||
"source_url": "https://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "[1912.01703v1] PyTorch: An Imperative Style, High-Performance Deep Learning Library",
|
||
"snippet": "[1912.01703v1] PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n# Title:PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n> Abstract:Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several common benchmarks.",
|
||
"source_url": "https://arxiv.org/abs/1912.01703v1",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1912.01703v1",
|
||
"_exa_published_date": null
|
||
},
|
||
{
|
||
"title": "PyTorch: An Imperative Style, High-Performance Deep Learning ...",
|
||
"snippet": "[1912.01703] PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n# PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\nAdam Paszke University of Warsaw adam.paszke@gmail.com Sam Gross Facebook AI Research sgross@fb.com Francisco Massa Facebook AI Research fmassa@fb.com Adam Lerer Facebook AI Research alerer@fb.com James Bradbury Google jekbradbury@gmail.com Gregory Chanan Facebook AI Research gchanan@fb.com Trevor Killeen Self Employed killeent@cs.washington.edu Zeming Lin Facebook AI Research zlin@fb.com Natalia Gimelshein NVIDIA ngimelshein@nvidia.com Luca Antiga Orobix luca.antiga@orobix.com Alban Desmaison Oxford University alban@robots.ox.ac.uk Andreas Köpf Xamla andreas.koepf@xamla.com Edward Yang Facebook AI Research ezyang@fb.com Zach DeVito Facebook AI Research zdevito@cs.stanford.edu Martin Raison Nabla martinraison@gmail.com Alykhan Tejani Twitter atejani@twitter.com Sasank Chilamkurthy Qure.ai sasankchilamkurthy@gmail.com Benoit Steiner Facebook AI Research benoitsteiner@fb.com Lu Fang Facebook lufang@fb.com Junjie Bai Facebook jbai@fb.com Soumith Chintala Facebook AI Research soumith@gmail.com\n[...]\nDeep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popu",
|
||
"source_url": "https://arxiv.org/abs/1912.01703",
|
||
"discovered_for": [
|
||
"method.training"
|
||
],
|
||
"_exa_id": "https://arxiv.org/abs/1912.01703",
|
||
"_exa_published_date": "2019-12-03T00:00:00.000Z"
|
||
}
|
||
],
|
||
"n_before_dedup": 90,
|
||
"n_after_dedup": 71,
|
||
"n_removed": 19
|
||
} |