feat(vla): add SmolVLA conditioning and experiment artifacts
This commit is contained in:
@@ -0,0 +1,904 @@
|
||||
{
|
||||
"candidates": [
|
||||
{
|
||||
"title": "Visuomotor policy learning via action diffusion - ACM Digital Library",
|
||||
"snippet": "Diffusion policy: : Visuomotor policy learning via action diffusion: International Journal of Robotics Research: Vol 4\n[...]\n, No 1\n[...]\nother-periodical;requested\n[...]\n:string:\n[...]\nPublication Websites;subPage:\n[...]\n:Basic Abstract\n[...]\n;page:string:Article/Chapter View;ctype:string:Journal Content;group\n[...]\n:acm-\n[...]\ntype>other-periodical;website:website:\n[...]\n-site;\n[...]\n:issue:\n[...]\n\\:10.5\n[...]\n55/rbrs.20\n[...]\n.44.issue-10-11;csubtype:string:Periodical;taxonomy:taxonomy:acm-pubtype;pageGroup:string:Publication Pages\"> skip to main content\n\n \n\n \n\nContents\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 15 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion mode",
|
||||
"source_url": "https://dl.acm.org/doi/10.1177/02783649241273668",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://dl.acm.org/doi/10.1177/02783649241273668",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Visuomotor Policy Learning via Action Diffusion - arXiv",
|
||||
"snippet": "# Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 15 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details is available diffusion-policy.cs.columbia.edu\n[...]\nIn this work, we s",
|
||||
"source_url": "https://arxiv.org/abs/2303.04137",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2303.04137",
|
||||
"_exa_published_date": "2023-03-07T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2303.04137v5] Diffusion Policy: Visuomotor Policy Learning via Action Diffusion",
|
||||
"snippet": "[2303.04137v5] Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\n# Title:Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\n> Abstract:This paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot's visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models.",
|
||||
"source_url": "http://arxiv.org/abs/2303.04137v5",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "http://arxiv.org/abs/2303.04137v5",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Visuomotor Policy Learning via Action Diffusion - arXiv",
|
||||
"snippet": "Diffusion Policy\n\n# Diffusion Policy\n\nCheng Chi1, Siyuan Feng2, Yilun Du3, Zhenjia Xu1, Eric Cousineau2, Benjamin Burchfiel2, Shuran Song1 1 Columbia University 2 Toyota Research Institute 3 MIT https://diffusion-policy.cs.columbia.edu\n\n# Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n\nCheng Chi1, Siyuan Feng2, Yilun Du3, Zhenjia Xu1, Eric Cousineau2, Benjamin Burchfiel2, Shuran Song1 1 Columbia University 2 Toyota Research Institute 3 MIT https://diffusion-policy.cs.columbia.edu\n\n###### Abstract\n\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper",
|
||||
"source_url": "https://arxiv.org/html/2502.12371v1",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2502.12371v1",
|
||||
"_exa_published_date": "2025-02-17T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2303.04137] Diffusion Policy - ar5iv - arXiv",
|
||||
"snippet": "# Diffusion Policy: Visuomotor Policy Learning via Action Diffusion\n[...]\nThis paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot’s visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details will be publicly available.\n[...]\nIn this work, we seek to address thi",
|
||||
"source_url": "https://ar5iv.labs.arxiv.org/html/2303.04137",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://ar5iv.labs.arxiv.org/html/2303.04137",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2006.11239v2] Denoising Diffusion Probabilistic Models",
|
||||
"snippet": "[2006.11239\n[...]\n] Denoising Diffusion Probabilistic Models\n[...]\n# Title:Denoising Diffusion Probabilistic Models\n[...]\nAuthors: Jonathan Ho, Ajay Jain, Pieter Abbeel\n[...]\n> Abstract:We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our implementation is available at this https URL\n\n \n\nhttps://doi.org/10.48550/arXiv.2006.11239\n\n \n\narXiv-issued DOI via DataCite",
|
||||
"source_url": "https://arxiv.org/abs/2006.11239v2",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2006.11239v2",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2006.11239] Denoising Diffusion Probabilistic Models - arXiv",
|
||||
"snippet": "[2006.11239] Denoising Diffusion Probabilistic Models\n[...]\n# Denoising Diffusion Probabilistic Models\n[...]\nJonathan Ho UC Berkeley jonathanho@berkeley.edu &Ajay Jain UC Berkeley ajayj@berkeley.edu &Pieter Abbeel UC Berkeley pabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our implementation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [53]. A diffusion probabilistic model (which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using variational inference to produce samples matching the data after finite time. Transitions of this chain are learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the data in the opposite direction of s",
|
||||
"source_url": "https://arxiv.org/abs/2006.11239",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2006.11239",
|
||||
"_exa_published_date": "2020-06-19T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Denoising Diffusion Probabilistic Models",
|
||||
"snippet": "Denoising Diffusion Probabilistic Models \n\nAuthorFeedback Bibtex MetaReview Paper Review Supplemental\n\n## Abstract\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN.\n\n \n\nDo not remove: This comment is monitored to verify that the site is working properly",
|
||||
"source_url": "https://papers.nips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://papers.nips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "Denoising Diffusion Probabilistic Models\n[...]\njonathanho@berkeley.edu\n[...]\nAjay Jain\n[...]\najayj@berkeley.edu\n[...]\nPieter Abbeel\n[...]\npabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models,\n[...]\na class of latent variable models inspired by considerations from nonequilibrium\n[...]\nthermodynamics. Our best results are obtained by training on a weighted variational\n[...]\nbound designed according to a novel connection between diffusion probabilistic\n[...]\nmodels and denoising score matching with Langevin dynamics, and our models nat\u0002urally admit a progressive lossy decompression scheme that can be interpreted as a\n[...]\ngeneralization of autoregressive decoding. On the unconditional CIFAR10 dataset,\n[...]\nwe obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On\n[...]\n256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our imple\u0002mentation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [50]. A diffusion probabilistic model\n[...]\n(which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using\n[...]\nvariational inference to produce samples matching the data after finite time. Transitions of this chain\n[...]\nare learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the\n[...]\ndata in the opposite direction of sampling until signal",
|
||||
"source_url": "https://arxiv.org/pdf/2006.11239v1",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2006.11239v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Denoising Diffusion Probabilistic Models - arXiv",
|
||||
"snippet": "Denoising Diffusion Probabilistic Models\n[...]\njonathanho@berkeley.edu\n[...]\nAjay Jain\n[...]\najayj@berkeley.edu\n[...]\nPieter Abbeel\n[...]\npabbeel@cs.berkeley.edu\n[...]\nWe present high quality image synthesis results using diffusion probabilistic models,\n[...]\na class of latent variable models inspired by considerations from nonequilibrium\n[...]\nthermodynamics. Our best results are obtained by training on a weighted variational\n[...]\nbound designed according to a novel connection between diffusion probabilistic\n[...]\nmodels and denoising score matching with Langevin dynamics, and our models nat\u0002urally admit a progressive lossy decompression scheme that can be interpreted as a\n[...]\ngeneralization of autoregressive decoding. On the unconditional CIFAR10 dataset,\n[...]\nwe obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On\n[...]\n256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our imple\u0002mentation is available at https://github.com/hojonathanho/diffusion.\n[...]\nThis paper presents progress in diffusion probabilistic models [53]. A diffusion probabilistic model\n[...]\n(which we will call a “diffusion model” for brevity) is a parameterized Markov chain trained using\n[...]\nvariational inference to produce samples matching the data after finite time. Transitions of this chain\n[...]\nare learned to reverse a diffusion process, which is a Markov chain that gradually adds noise to the\n[...]\ndata in the opposite direction of sampling until signal",
|
||||
"source_url": "https://arxiv.org/pdf/2006.11239",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2006.11239",
|
||||
"_exa_published_date": "2020-12-16T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2212.09748] Scalable Diffusion Models with Transformers - arXiv",
|
||||
"snippet": "William Peebles* UC Berkeley Saining Xie New York University\n[...]\nWe explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops—through increased transformer depth/width or increased number of input tokens—consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512 $\\times$ 512 and 256 $\\times$ 256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.\n[...]\n, or Di\n[...]\nfor short. Di\n[...]\nadhere to the best practices of Vision Transformers (ViTs) [10], which have been shown to scale more effectively for visual recognition than\n[...]\nconvolutional networks (e.g., Res\n[...]\n[15]).\n[...]\nMore specifically, we study the scaling behavior of transformers with respect to network complexity vs. sample quality. We show that by constructing and benchmarking the DiT design space under the Latent Diffusion Models (LDMs) [48] framework, where diffusion models are trained within a VAE’s latent space, we can successfully replace the U-Net backbone with a transformer. We further show that DiTs are scalable architectures for diff",
|
||||
"source_url": "https://arxiv.org/abs/2212.09748",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2212.09748",
|
||||
"_exa_published_date": "2022-12-19T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore",
|
||||
"snippet": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore\n\n \n\n \n\n \n\n### IEEE Account",
|
||||
"source_url": "https://ieeexplore.ieee.org/document/10377858/",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://ieeexplore.ieee.org/document/10377858/",
|
||||
"_exa_published_date": "2025-05-14T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2212.09748v1] Scalable Diffusion Models with Transformers",
|
||||
"snippet": "[2212.09748v1] Scalable Diffusion Models with Transformers\n[...]\n# Title:Scalable Diffusion Models with Transformers\n[...]\nAuthors: William Peebles, Saining Xie\n[...]\n> Abstract:We explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops -- through increased transformer depth/width or increased number of input tokens -- consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512x512 and 256x256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.",
|
||||
"source_url": "http://arxiv.org/abs/2212.09748v1",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "http://arxiv.org/abs/2212.09748v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Scalable Diffusion Models with Transformers | Semantic Scholar",
|
||||
"snippet": "[PDF] Scalable Diffusion Models with Transformers | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Peebles2022ScalableDM,\n title={Scalable Diffusion Models with Transformers},\n author={William S. Peebles and Saining Xie},\n journal={2023 IEEE/CVF International Conference on Computer Vision (ICCV)},\n year={2022},\n pages={4172-4182},\n url={https://api.semanticscholar.org/CorpusID:254854389}\n}\n```",
|
||||
"source_url": "https://www.semanticscholar.org/reader/736973165f98105fec3729b7db414ae4d80fcbeb",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/reader/736973165f98105fec3729b7db414ae4d80fcbeb",
|
||||
"_exa_published_date": "2022-12-19T14:39:50.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Scalable Diffusion Models with Transformers - IEEE Xplore",
|
||||
"snippet": "Scalable Diffusion Models with Transformers | IEEE Conference Publication | IEEE Xplore\n**\n\n### IEEE Account\n* Change Username/Password\n* Update Address\n### Purchase Details\n* Payment Options\n* Order History\n* View Purchased Documents\n### Profile Information\n* Communications Preferences\n* Profession and Education\n* Technical Interests\n### Need Help?\n* **US & Canada:**+1 800 678 4333\n* **Worldwide:**+1 732 981 0060\n* Contact & Support\n* About IEEE*Xplore*\n* Contact Us\n* Help\n* Accessibility\n* Terms of Use\n* Nondiscrimination Policy\n* Sitemap\n* Privacy & Opting Out of Cookies\nA not-for-profit organization, IEEE is the world's largest technical professional organization dedicated to advancing technology for the benefit of humanity.\n© Copyright 2025 IEEE - All rights reserved. Use of this web site signifies your agreement to the terms and conditions.\n**",
|
||||
"source_url": "https://ieeexplore.ieee.org/iel7/10376473/10376477/10377858.pdf",
|
||||
"discovered_for": [
|
||||
"rw.diffusion_policy"
|
||||
],
|
||||
"_exa_id": "https://ieeexplore.ieee.org/iel7/10376473/10376477/10377858.pdf",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Flow Matching for Generative Modeling - arXiv",
|
||||
"snippet": "FLOW MATCHING FOR GENERATIVE MODELING\n[...]\nYaron Lipman1,2 Ricky T. Q. Chen1 Heli Ben-Hamu2 Maximilian Nickel1 Matt Le1\n[...]\nWe introduce a new paradigm for generative modeling built on Continuous\n[...]\n(CNFs), allowing us to train CNFs at unprecedented scale.\n[...]\nSpecifically, we present the notion of Flow Matching (FM), a simulation-free\n[...]\napproach for training CNFs based on regressing vector fields of fixed conditional\n[...]\nprobability paths. Flow Matching is compatible with a general family of Gaussian\n[...]\nprobability paths for transforming between noise and data samples—which\n[...]\nsubsumes existing diffusion paths as specific instances. Interestingly, we find\n[...]\nthat employing FM with diffusion paths results in a more robust and stable\n[...]\nalternative for training diffusion models. Furthermore, Flow Matching opens\n[...]\nthe door to training CNFs with other, non-diffusion probability paths. An\n[...]\ninstance of particular interest is using Optimal Transport (OT) displacement\n[...]\ninterpolation to define the conditional probability paths. These paths are more\n[...]\nefficient than diffusion paths, provide faster training and sampling, and result in\n[...]\nbetter generalization. Training CNFs using Flow Matching on ImageNet leads\n[...]\nto consistently better performance than alternative diffusion-based methods in\n[...]\nterms of both likelihood and sample quality, and allows fast and reliable sample\n[...]\ngeneration using off-the-shelf numerical ODE solvers.\n",
|
||||
"source_url": "https://arxiv.org/pdf/2210.02747",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2210.02747",
|
||||
"_exa_published_date": "2023-02-08T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "FLOW MATCHING FOR GENERATIVE MODELING\n[...]\nYaron Lipman1,2 Ricky T. Q. Chen1 Heli Ben-Hamu2 Maximilian Nickel1 Matt Le1\n[...]\nWe introduce a new paradigm for generative modeling built on Continuous\n[...]\nNormalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale.\n[...]\nSpecifically, we present the notion of Flow Matching (FM), a simulation-free\n[...]\napproach for training CNFs based on regressing vector fields of fixed conditional\n[...]\nprobability paths. Flow Matching is compatible with a general family of Gaussian\n[...]\nprobability paths for transforming between noise and data samples—which\n[...]\nsubsumes existing diffusion paths as specific instances. Interestingly, we find\n[...]\nthat employing FM with diffusion paths results in a more robust and stable\n[...]\nalternative for training diffusion models. Furthermore, Flow Matching opens\n[...]\nthe door to training CNFs with other, non-diffusion probability paths. An\n[...]\ninstance of particular interest is using Optimal Transport (OT) displacement\n[...]\ninterpolation to define the conditional probability paths. These paths are more\n[...]\nefficient than diffusion paths, provide faster training and sampling, and result in\n[...]\nbetter generalization. Training CNFs using Flow Matching on ImageNet leads\n[...]\nto consistently better performance than alternative diffusion-based methods in\n[...]\nterms of both likelihood and sample quality, and allows fast and reliable sample\n[...]\ngeneration using off-the-shelf numer",
|
||||
"source_url": "https://openreview.net/pdf?id=PqvMRDCJT9t",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/pdf?id=PqvMRDCJT9t",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2210.02747] Flow Matching for Generative Modeling - arXiv",
|
||||
"snippet": "[221\n[...]\n] Flow Matching for Generative Modeling\n[...]\n# Flow Matching for Generative Modeling\n[...]\nYaron Lipman1,2 Ricky T. Q. Chen1 Heli Ben-Hamu2 Maximilian Nickel1 Matt Le1 1Meta AI (FAIR) 2Weizmann Institute of Science\n[...]\nWe introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths. Flow Matching is compatible with a general family of Gaussian probability paths for transforming between noise and data samples—which subsumes existing diffusion paths as specific instances. Interestingly, we find that employing FM with diffusion paths results in a more robust and stable alternative for training diffusion models. Furthermore, Flow Matching opens the door to training CNFs with other, non-diffusion probability paths. An instance of particular interest is using Optimal Transport (OT) displacement interpolation to define the conditional probability paths. These paths are more efficient than diffusion paths, provide faster training and sampling, and result in better generalization. Training CNFs using Flow Matching on ImageNet leads to consistently better performance than alternative diffusion-based methods in terms of both likelihood and sample quality, and allows fast and reliable sample generation using off-the-s",
|
||||
"source_url": "https://arxiv.org/abs/2210.02747",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2210.02747",
|
||||
"_exa_published_date": "2022-10-06T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Flow Matching for Generative Modeling | Semantic Scholar",
|
||||
"snippet": "```\n@article{Lipman2022FlowMF,\n title={Flow Matching for Generative Modeling},\n author={Yaron Lipman and Ricky T. Q. Chen and Heli Ben-Hamu and Maximilian Nickel and Matt Le},\n journal={ArXiv},\n year={2022},\n volume={abs/2210.02747},\n url={https://api.semanticscholar.org/CorpusID:252734897}\n}\n```",
|
||||
"source_url": "https://www.semanticscholar.org/reader/af68f10ab5078bfc519caae377c90ee6d9c504e9",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/reader/af68f10ab5078bfc519caae377c90ee6d9c504e9",
|
||||
"_exa_published_date": "2022-10-06T03:39:18.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Flow Matching for Generative Modeling - OpenReview",
|
||||
"snippet": "Flow Matching for Generative Modeling | OpenReview\n\n## Flow Matching for Generative Modeling\n\nICLR 2023 notable top 25%Readers: Everyone\n\nKeywords: continuous normalizing flows, generative models\n\nAbstract: We introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths. Flow Matching is compatible with a general family of Gaussian probability paths for transforming between noise and data samples---which subsumes existing diffusion paths as specific instances. Interestingly, we find that employing FM with diffusion paths results in a more robust and stable alternative for training diffusion models. Furthermore, Flow Matching opens the door to training CNFs with other, non-diffusion probability paths. An instance of particular interest is using Optimal Transport (OT) displacement interpolation to define the conditional probability paths. These paths are more efficient than diffusion paths, provide faster training and sampling, and result in better generalization. Training CNFs using Flow Matching on ImageNet leads to consistently better performance than alternative diffusion-based methods in terms of both likelihood and sample quality, and allows fast and reliable sample generation using off-the-shelf numerical ODE solve",
|
||||
"source_url": "https://openreview.net/forum?id=PqvMRDCJT9t",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=PqvMRDCJT9t",
|
||||
"_exa_published_date": "2022-09-29T08:12:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2505.13447] Mean Flows for One-step Generative Modeling - arXiv",
|
||||
"snippet": "# Mean Flows for One-step Generative Modeling\n[...]\nZhengyang Geng1 Mingyang Deng2 Xingjian Bai2 J. Zico Kolter1 Kaiming He2 1CMU 2MIT Work partly done when visiting MIT.\n[...]\nWe propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n[...]\nIn this work, we propose a principled and effective framework, termed MeanFlow, for one-step generation. The core idea is to introduce a new ground-truth field representing the average velocity, in contrast to the instantaneous velocity typically modeled in Flow Matching. Average velocity is defined as the ratio of displacement to a time interval, with displacement give",
|
||||
"source_url": "https://arxiv.org/abs/2505.13447",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2505.13447",
|
||||
"_exa_published_date": "2025-05-19T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Mean Flows for One-step Generative Modeling",
|
||||
"snippet": "# Mean Flows for One-step Generative Modeling\n[...]\nZhengyang Geng1 Mingyang Deng2 Xingjian Bai2 J. Zico Kolter1 Kaiming He2 1CMU 2MIT Work partly done when visiting MIT.\n[...]\nWe propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n[...]\nIn this work, we propose a principled and effective framework, termed MeanFlow, for one-step generation. The core idea is to introduce a new ground-truth field representing the average velocity, in contrast to the instantaneous velocity typically modeled in Flow Matching. Average velocity is defined as the ratio of displacement to a time interval, with displacement give",
|
||||
"source_url": "https://arxiv.org/html/2505.13447",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2505.13447",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "Mean Flows for One-step Generative Modeling\n[...]\nZhengyang Geng1∗ Mingyang Deng2 Xingjian Bai2 J. Zico Kolter1 Kaiming He2\n[...]\nWe propose a principled and effective framework for one-step generative modeling.\n[...]\nWe introduce the notion of average velocity to characterize flow fields, in contrast to\n[...]\ninstantaneous velocity modeled by Flow Matching methods. A well-defined identity\n[...]\nbetween average and instantaneous velocities is derived and used to guide neural\n[...]\nnetwork training. Our method, termed the MeanFlow model, is self-contained and\n[...]\nrequires no pre-training, distillation, or curriculum learning. MeanFlow demon\u0002strates strong empirical performance: it achieves an FID of 3.43 with a single\n[...]\nfunction evaluation (1-NFE) on ImageNet 256×256 trained from scratch, signifi\u0002cantly outperforming previous state-of-the-art one-step diffusion/flow models. Our\n[...]\nstudy substantially narrows the gap between one-step diffusion/flow models and\n[...]\ntheir multi-step predecessors, and we hope it will motivate future research to revisit\n[...]\nIn this work, we propose a principled and effective framework, termed MeanFlow, for one-step\n[...]\ngeneration. The core idea is to introduce a new ground-truth field representing the average velocity,\n[...]\nin contrast to the instantaneous velocity typically modeled in Flow Matching. Average velocity is\n[...]\ndefined as the ratio of displacement to a time interval, with displacement given by the time integral of\n[...",
|
||||
"source_url": "https://openreview.net/pdf/611127035c61b58c8523c81094b8109e0d8f84fc.pdf",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/pdf/611127035c61b58c8523c81094b8109e0d8f84fc.pdf",
|
||||
"_exa_published_date": "2025-10-24T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Mean Flows for One-step Generative Modeling - OpenReview",
|
||||
"snippet": "Mean Flows for One-step Generative Modeling | OpenReview\n\n## Mean Flows for One-step Generative Modeling\n\n### Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, Kaiming He\n\nNeurIPS 2025 oralEveryone Revisions BibTeX CC BY 4.0\n\nKeywords: Generative Models\n\nAbstract: We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the \\textit{MeanFlow} model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256$\\times$256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.\n\nPrimary Area: Deep learning (e.g., architectures, generative models, optimization for deep networks, foundation models, LLMs)\n\nSubmission Number: 754\n\nLoading",
|
||||
"source_url": "https://openreview.net/forum?id=uWj4s7rMnR",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=uWj4s7rMnR",
|
||||
"_exa_published_date": "2025-10-29T14:53:17.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Mean Flows for One-step Generative Modeling | Semantic Scholar",
|
||||
"snippet": "[PDF] Mean Flows for One-step Generative Modeling | Semantic Scholar\n[...]\n```\n@article{Geng2025MeanFF,\n title={Mean Flows for One-step Generative Modeling},\n author={Zhengyang Geng and Mingyang Deng and Xingjian Bai and J. Zico Kolter and Kaiming He},\n journal={ArXiv},\n year={2025},\n volume={abs/2505.13447},\n url={https://api.semanticscholar.org/CorpusID:278769814}\n}\n```",
|
||||
"source_url": "https://www.semanticscholar.org/reader/19df654b0d0f634a451564346a09af8bd348dac0",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/reader/19df654b0d0f634a451564346a09af8bd348dac0",
|
||||
"_exa_published_date": "2025-05-19T06:36:24.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Improved Mean Flows: On the Challenges of Fastforward Generative ...",
|
||||
"snippet": "# Improved Mean Flows: On the Challenges of Fastforward Generative Models\n[...]\nZhengyang Geng1,2,3, Yiyang Lu4,2,∗ Zongze Wu3 Eli Shechtman3 J. Zico Kolter1 Kaiming He2 1CMU 2MIT 3Adobe 4THU Equal contribution. Part of this work was done when Z. Geng was interning at Adobe and MIT, and when Y. Lu was interning at MIT.\n[...]\nMeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its “fastforward” nature introduces key challenges in both the training objective and the guidance mechanism. First, the original MF’s training target depends not only on the underlying ground-truth fields but also on the network itself. To address this issue, we recast the objective as a loss on the instantaneous velocity $v$ , re-parameterized by a network that predicts the average velocity $u$ . Our reformulation yields a more standard regression problem and improves the training stability. Second, the original MF fixes the classifier-free guidance scale during training, which sacrifices flexibility. We tackle this issue by formulating guidance as explicit conditioning variables, thereby retaining flexibility at test time. The diverse conditions are processed through in-context conditioning, which reduces model size and benefits performance. Overall, our improved MeanFlow (iMF) method, trained entirely from scratch, achieves 1.72 FID with a single function evaluation (1-NFE) on ImageNet 256 $\\times$ 256. iMF substantially outperforms prior methods of t",
|
||||
"source_url": "https://arxiv.org/abs/2512.02012",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2512.02012",
|
||||
"_exa_published_date": "2025-12-01T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "probability path. At the same time, the design space around this training objective is rapidly expand\u0002ing with emerging techniques. For example, Mean Flow (Geng et al., 2025a) replaces instantaneous\n[...]\nvelocities with average velocities to enable strong one-step generation, improved Mean Flow vari\u0002ants (Geng et al., 2025b) reformulate the objective and guidance mechanism for better stability and\n[...]\nflexibility,\n[...]\n” (Li\n[...]\npatch Transformers.\n[...]\ncurriculum sampling, a\n[...]\n-phase curriculum that\n[...]\n(Logit-Normal)\n[...]\na coverage-focused distribution (Uniform). This method outperforms static base\u0002lines, achieving a 1\n[...]\n% relative improvement in FID over standard Uniform sampling\n[...]\nrevisits a central but often under-specified design choice in flow-based generative model\u0002ing: the timestep sampling distribution\n[...]\n). We identify a performance paradox in standard training:\n[...]\nMotivated by this analysis, we propose a two-phase curriculum for timestep sampling. The cur\u0002riculum uses a Logit-Normal distribution in a structure-learning phase to exploit fast early progress,\n[...]\nthen switches to Uniform sampling in a refinement phase to restore coverage of\n[...]\n-10, this simple schedule improves the best FID from 3.85 (Uniform) to 3.22 (a 16.4%\n[...]\nZhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. Mean flows for\n[...]\none-step generative modeling. NeurIPS, 2025a.\n[...]\nZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman, J.",
|
||||
"source_url": "https://arxiv.org/pdf/2603.12517",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2603.12517",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "Tianhong Li 1 Zhengyang Geng 2 Kaiming He 1\n[...]\nModern diffusion/flow-based models for image\n[...]\nmade encouraging progress on each aspect in\u0002dividually, paving the way toward one-step dif\u0002fusion/flow without latents. In this work, we\n[...]\n“pixel MeanFlow” (pMF). Our core guideline is\n[...]\nformulate the network output space and the\n[...]\nloss space separately. The network target is de\u0002signed to be on a presumed low-dimensional im\u0002age manifold (i.e., x-prediction), while the loss\n[...]\nis defined via MeanFlow in the velocity space.\n[...]\nWe introduce a simple transformation between\n[...]\nthe image manifold and the average velocity field.\n[...]\nIn experiments, pMF achieves strong results for\n[...]\none-step latent-free generation on ImageNet at\n[...]\n256×256 resolution (2.22 FID) and 512×512 res\u0002olution (2.48 FID), filling a key missing piece in\n[...]\nIn this work, we propose pixel MeanFlow (pMF) for one\u0002step latent-free image generation. pMF follows the im\u0002proved MeanFlow (iMF) (Geng et al., 2025b) that learns\n[...]\nthe average velocity field (namely, u) using a loss defined\n[...]\nin the space of instantaneous velocity (namely, v). On the\n[...]\nother hand, following JiT (Li & He, 2025), pMF directly\n[...]\nparameterizes a denoised-image-like quantity (namely, x\u0002prediction), which is expected to lie on a low-dimensional\n[...]\nmanifold. To accommodate both formulations, we intro\u0002duce a conversion that relates the fields v, u, and x. We\n[...]\nempirically show that this formula",
|
||||
"source_url": "https://arxiv.org/pdf/2601.22158",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2601.22158",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "In this work, we propose Discrete MeanFlow (DMF) training curriculum, a budget-friendly frame\u0002work designed to bridge the gap between standard flow models and the MeanFlow identity for fast\n[...]\n. Our approach\n[...]\na staged curriculum\n[...]\nZhengyang Geng, Mingyang Deng, Xingjian Bai, J. Zico Kolter, and Kaiming He. Mean flows for\n[...]\n, 20\n[...]\n. URL https\n[...]\narxiv.org/\n[...]\n/2505.13447.\n[...]\nZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman, J. Zico Kolter, and Kaiming He. Im\u0002proved mean flows: On the challenges of fastforward generative models, 2025b. URL https:\n[...]\n//arxiv.org/abs/2512.02012.",
|
||||
"source_url": "https://arxiv.org/pdf/2604.08837",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2604.08837",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2511.19065] Understanding, Accelerating, and Improving MeanFlow Training",
|
||||
"snippet": "Accelerating, and Improving MeanFlow Training\n[...]\n# Title:Understanding, Accelerating, and Improving MeanFlow Training\n\nAuthors: Jin-Young Kim, Hyojun Go, Lea Bogensperger, Julius Erbach, Nikolai Kalischek, Federico Tombari, Konrad Schindler, Dominik Narnhofer\n\nView PDF HTML (experimental)\n[...]\n> Abstract:MeanFlow promises high-quality generative modeling in few steps, by jointly learning instantaneous and average velocity fields. Yet, the underlying training dynamics remain unclear. We analyze the interaction between the two velocities and find: (i) well-established instantaneous velocity is a prerequisite for learning average velocity; (ii) learning of instantaneous velocity benefits from average velocity when the temporal gap is small, but degrades as the gap increases; and (iii) task-affinity analysis indicates that smooth learning of large-gap average velocities, essential for one-step generation, depends on the prior formation of accurate instantaneous and small-gap average velocities. Guided by these observations, we design an effective training scheme that accelerates the formation of instantaneous velocity, then shifts emphasis from short- to long-interval average velocity. Our enhanced MeanFlow training yields faster convergence and significantly better few-step generation: With the same DiT-XL backbone, our method reaches an impressive FID of 2.87 on 1-NFE ImageNet 256x256, compared to 3.43 for the conventional MeanFlow baseline. Alternatively, our method matche",
|
||||
"source_url": "https://arxiv.org/abs/2511.19065",
|
||||
"discovered_for": [
|
||||
"rw.mean_flow"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2511.19065",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "RT-1: Robotics Transformer for Real-World Control at Scale",
|
||||
"snippet": "By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project’s website and videos can be found at robotics-transformer1.github.io\n[...]\nThe second challenge lies in the design of the model itself. Effective robotic multi-task learning requires a high capacity model, and Transformer (Vaswani et al., 2017) models excel in this regard, particularly when it is necessary to learn many tasks conditioned, as in our case, on language instruct",
|
||||
"source_url": "https://arxiv.org/html/2212.06817",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2212.06817",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "RT-1: Robotics Transformer for Real-World Control at Scale - arXiv",
|
||||
"snippet": "By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project’s website and videos can be found at robotics-transformer1.github.io\n[...]\nThe second challenge lies in the design of the model itself. Effective robotic multi-task learning requires a high capacity model, and Transformer (Vaswani et al., 2017) models excel in this regard, particularly when it is necessary to learn many tasks conditioned, as in our case, on language instruct",
|
||||
"source_url": "https://arxiv.org/abs/2212.06817",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2212.06817",
|
||||
"_exa_published_date": "2022-12-13T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "RT-1: Robotics Transformer for Real-World Control at Scale - arXiv",
|
||||
"snippet": "By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project’s website and videos can be found at robotics-transformer1.github.io\n[...]\nThe second challenge lies in the design of the model itself. Effective robotic multi-task learning requires a high capacity model, and Transformer (Vaswani et al., 2017) models excel in this regard, particularly when it is necessary to learn many tasks conditioned, as in our case, on language instruct",
|
||||
"source_url": "https://arxiv.org/html/2212.06817v2",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2212.06817v2",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Bringing the RT-1-X Foundation Model to a SCARA robot",
|
||||
"snippet": "Traditional robotic systems require specific training data for each task, environment, and robot form. While recent advancements in machine learning have enabled models to generalize across new tasks and environments, the challenge of adapting these models to entirely new settings remains largely unexplored. This study addresses this by investigating the generalization capabilities of the RT-1-X robotic foundation model to a type of robot unseen during its training: a SCARA robot from UMI-RTX.\n[...]\nRecent breakthroughs in machine learning and artificial intelligence suggest that training on large, diverse datasets can lead to highly\n[...]\nwhich often exceed\n[...]\nperformance of models developed for\n[...]\ndatasets tailored to\n[...]\nAs a result,\n[...]\nfield has been exploring more\n[...]\ncan adapt to\n[...]\n. Recent advancements like transformer\n[...]\n’s RT-1 [brohan_rt-1_2022], which demonstrate the potential for\n[...]\nGoogle’s RT-1 model is an impressive work, tested on a collection of real-world robotic experiences, where in different institutes a fleet of robots were performing 700 tasks [brohan_rt-1_2022]. The robots in the training set, such as the Franka, Kuka iiwa, UR5 and the EveryDay robot, can move their end-effector in a spherical working-space. None of the robots in the dataset is of the SCARA (Selective Compliance Assembly Robot Arm) type. With a SCARA robot the movement of z-axis is decoupled from the movement in the x-y plane, which gives a SCARA robot an kidney ",
|
||||
"source_url": "https://arxiv.org/html/2409.03299v1",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2409.03299v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "generalization capabilities of\n[...]\nfoundation models have\n[...]\n’s RT-1\n[...]\nGoogle’s RT-1 model is an impressive work, tested on a collection of real\u0002\n[...]\nrobotic experiences, where in\n[...]\ninstitutes a fleet of robots were\n[...]\nstudy is to see if generalization capabilities of the Google’s RT-1 model can\n[...]\nextended to an unseen robot\n[...]\nof a complete different type.\n[...]\nThe RT-1 model [2] was presented as a joint effort between Robotics at Google,\n[...]\nEveryday Robots, and Google Research, at the end of 2022. The purpose of RT\u00021 is to investigate if it is possible to train a single, capable, multi-task model\n[...]\non data consisting of a wide variety of robotic tasks, and to find out if such\n[...]\na model brings the same benefits observed in other domains, namely zero-shot\n[...]\ngeneralization to new tasks, environments, and objects.\n[...]\nWith RT-1, the authors present a model architecture along with a signif\u0002icant training dataset that fulfills those requirements, as well as demonstrate\n[...]\nof this model [2\n[...]\nBecause initial experiments [17] showed that zero-shot generalization was not\n[...]\npossible to the unseen UMI robotic embodiment, the decision was made to fine\u0002tune the model.\n[...]\nevaluate the generalization potential of\n[...]\n-1-\n[...]\nThis study explored the generalization capabilities of the RT-1-X model, particu\u0002larly its ability to adapt to a robot type not seen before. A dataset of demonstra\u0002tions was collected using the UMI robot, and",
|
||||
"source_url": "https://arxiv.org/pdf/2409.03299",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2409.03299",
|
||||
"_exa_published_date": "2024-09-06T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | OpenReview",
|
||||
"snippet": "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | OpenReview\n[...]\n## RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nTL;DR: Vision-language models, trained on Internet-scale data, can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning.\n[...]\nAbstract: We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows th",
|
||||
"source_url": "https://openreview.net/forum?id=XMQgwiJ7KSX",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=XMQgwiJ7KSX",
|
||||
"_exa_published_date": "2023-08-30T15:38:01.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] RT-2: Vision-Language-Action Models Transfer Web Knowledge to ...",
|
||||
"snippet": "Transfer Web Knowledge\n[...]\n# RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nWe study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot t",
|
||||
"source_url": "https://arxiv.org/pdf/2307.15818",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2307.15818",
|
||||
"_exa_published_date": "2023-07-28T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2307.15818] RT-2: Vision-Language-Action Models Transfer Web ...",
|
||||
"snippet": "Transfer Web Knowledge\n[...]\n# RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control\n[...]\nWe study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot t",
|
||||
"source_url": "https://arxiv.org/abs/2307.15818",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2307.15818",
|
||||
"_exa_published_date": "2023-07-28T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "language and point clouds and varying sequence lengths (from ∼ 200 to few thousand). Left: A 5B vision-language-action model from\n[...]\nthe RT-2 class [1] (sequence length: L = 196). The manipulation policy is conditioned on the text instruction. Right: A Point Cloud\n[...]\n-parameter vision-language-action models or\n[...]\nAs), into\n[...]\nspeeding up: (a) the class of recently introduced RT-2 models\n[...]\n[1], the first VLA robotic policies pre-trained on internet\u0002scale data, as well as (b) Point Cloud Transformer (PCT)\n[...]\n[12]), multi-modal sensor fusion [16], finally the first vision\u0002language-action robotic manipulation powered by massive\n[...]\nvision language models [1].\n[...]\nparameter models such as RT-2 [1].\n[...]\nto convert pre-trained or already fine-tuned Transformer\u0002based robotic policies of quadratic space and time complex\u0002ity (including massive billion-parameter vision-language\u0002action models or VLAs), into their efficient linear-attention\n[...]\neffectiveness of SARA-RT by speeding up: (a) the class\n[...]\nof the aforementioned RT-2 models, the first VLA robotic\n[...]\npolicies pre-trained on internet-scale data, as well as (b)\n[...]\nB. RT-2 vision-language-action models\n[...]\n1) The setting: Next we consider the class of RT-2 archi\u0002tectures from [1]. Those apply PaLI-X [36]) vision-language\u0002model (VLM) backbones to encode policies taking vision\n[...]\ninput and conditioned on natural language instructions. We\n[...]\nfocus on the 5B PaLI-X variant, as more practical ",
|
||||
"source_url": "https://arxiv.org/pdf/2312.01990",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2312.01990",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "robotics applications is particularly\n[...]\n. One notable\n[...]\nexample is RT-2, a system capable of generating low-level actions\n[...]\nrepresented in textual format from a given instruction alongside\n[...]\na sequence of historical actions and image observations. To\n[...]\nstimulate further research in this domain, we introduce an\n[...]\nopen-source implementation tailored for utilizing VLMs in\n[...]\ninstruction-based robot control. This implementation supports\n[...]\na variety of VLM architectures and facilitates straightforward\n[...]\nintegration of new models. We use our framework to train\n[...]\nmultiple VLMs and evaluate them on a physical robot. The\n[...]\nresults validate the practical efficacy of our framework, thus\n[...]\npaving the way for enhanced understanding and capabilities in\n[...]\ninstruction-based robot control systems. The code is available\n[...]\nA notable endeavor in this domain is RT-2 [12], which\n[...]\nemploys a VLM to interpret an instruction, past actions\n[...]\nand image observations, and generate subsequent actions in\n[...]\ntextual form. RT-2 has demonstrated the potential of creating\n[...]\ninstruction-based low-level robot control policies, showcasing\n[...]\nremarkable performance and notable generalization capabil\u0002ities. However, RT-2’s unavailability to the public and its\n[...]\nconsiderable scale, with 55 billion parameters, pose challenges\n[...]\nfor academic endeavors to engage with it effectively, hindering\n[...]\nfurther exploration and refinement of thi",
|
||||
"source_url": "https://openreview.net/pdf?id=nfm2qcV1S4",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/pdf?id=nfm2qcV1S4",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "OpenVLA: An Open-Source Vision-Language-Action Model - arXiv",
|
||||
"snippet": "OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n# OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\nLarge policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challenges, we introduce OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations. OpenVLA builds on a Llama 2 language model combined with a visual encoder that fuses pretrained features from DINOv2 and SigLIP. As a product of the added data diversity and new model components, OpenVLA demonstrates strong results for generalist manipulation, outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters. We further show that we can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving multiple objects and strong language",
|
||||
"source_url": "https://arxiv.org/abs/2406.09246",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2406.09246",
|
||||
"_exa_published_date": "2024-06-13T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] OpenVLA: An Open-Source Vision-Language-Action Model | Semantic Scholar",
|
||||
"snippet": "[PDF] OpenVLA: An Open-Source Vision-Language-Action Model | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Kim2024OpenVLAAO,\n title={OpenVLA: An Open-Source Vision-Language-Action Model},\n author={Moo Jin Kim and Karl Pertsch and Siddharth Karamcheti and Ted Xiao and Ashwin Balakrishna and Suraj Nair and Rafael Rafailov and Ethan Paul Foster and Grace Lam and Pannag R. Sanketi and Quan Vuong and Thomas Kollar and Benjamin Burchfiel and Russ Tedrake and Dorsa Sadigh and Sergey Levine and Percy Liang and Chelsea Finn},\n journal={ArXiv},\n year={2024},\n volume={abs/2406.09246},\n url={https://api.semanticscholar.org/CorpusID:270440391}\n}\n```",
|
||||
"source_url": "https://www.semanticscholar.org/reader/8f9ceb5ffad8e7a066dfc9d9aaa5153b714740ee",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/reader/8f9ceb5ffad8e7a066dfc9d9aaa5153b714740ee",
|
||||
"_exa_published_date": "2024-06-13T09:02:12.000Z"
|
||||
},
|
||||
{
|
||||
"title": "OpenVLA: An Open-Source Vision-Language-Action Model",
|
||||
"snippet": "OpenVLA: An Open-Source Vision-Language-Action Model | OpenReview\n\n## OpenVLA: An Open-Source Vision-Language-Action Model\n\n### Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, Chelsea Finn \n\nCoRL 2024everyonesince 05 Sept 2024\">Everyone Revisions BibTeX CC BY 4.0\n\nKeywords: Vision-Language-Action Models, Generalist Policies, Large-scale Robot Learning, Robotic Manipulation, Robotics, Vision-Language Models\n\nTL;DR: We introduce OpenVLA, a state-of-the-art, open-source 7B-parameter VLA model that obtains strong performance for cross-embodiment robot control out-of-the-box and can be easily adapted to new robot setups via parameter-efficient fine-tuning.\n\nAbstract: Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challeng",
|
||||
"source_url": "https://openreview.net/forum?id=ZMnD6QZAE6",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=ZMnD6QZAE6",
|
||||
"_exa_published_date": "2024-09-05T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[2406.09246v3] OpenVLA: An Open-Source Vision-Language-Action Model",
|
||||
"snippet": "[2406.09246v3] OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n# Title:OpenVLA: An Open-Source Vision-Language-Action Model\n[...]\n> Abstract:Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challenges, we introduce OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations. OpenVLA builds on a Llama 2 language model combined with a visual encoder that fuses pretrained features from DINOv2 and SigLIP. As a product of the added data diversity and new model components, OpenVLA demonstrates strong results for generalist manipulation, outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters. We further show that we can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving mult",
|
||||
"source_url": "https://arxiv.org/abs/2406.09246v3",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2406.09246v3",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Revision History for OpenVLA: An Open-Source... - OpenReview",
|
||||
"snippet": "Revisions | OpenReview\n\nLoading",
|
||||
"source_url": "https://openreview.net/revisions?id=ZMnD6QZAE6",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/revisions?id=ZMnD6QZAE6",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "A Vision-Language-Action Model for Affordable and Efficient Robotics",
|
||||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||||
"source_url": "https://arxiv.org/abs/2506.01844",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2506.01844",
|
||||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable ... - arXiv",
|
||||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||||
"source_url": "https://arxiv.org/pdf/2506.01844",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2506.01844",
|
||||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "SmolVLA: A vision-language-action model for affordable and ... - arXiv",
|
||||
"snippet": "SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\n# SmolVLA: A vision-language-action model for affordable and efficient robotics\n[...]\nVision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive–often with billions of parameters–leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10 $\\times$ larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic bench",
|
||||
"source_url": "https://arxiv.org/html/2506.01844v1",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2506.01844v1",
|
||||
"_exa_published_date": "2025-06-02T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics | Semantic Scholar",
|
||||
"snippet": "[PDF] SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics | Semantic Scholar \n\nNavigate Paper Download (opens in a new tab) Share\n[...]\n```\n@article{Shukor2025SmolVLAAV,\n title={SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics},\n author={Mustafa Shukor and Dana Aubakirova and Francesco Capuano and Pepijn Kooijmans and Steven Palma and Adil Zouitine and Michel Aractingi and Caroline Pascal and Martino Russi and Andr{\\'e}s Marafioti and Simon Alibert and Matthieu Cord and Thomas Wolf and R{\\'e}mi Cad{\\`e}ne},\n journal={ArXiv},\n year={2025},\n volume={abs/2506.01844},\n url={https://api.semanticscholar.org/CorpusID:279119427}\n}\n```",
|
||||
"source_url": "https://www.semanticscholar.org/reader/6ab4d113676d00e74b55e918fee4c7affaa8652f",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/reader/6ab4d113676d00e74b55e918fee4c7affaa8652f",
|
||||
"_exa_published_date": "2025-06-02T13:54:14.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots",
|
||||
"snippet": "By leveraging NF4 quantization and the llama-cpp runtime, the proposed LiteVLA implementation pioneers the CPU-only deployment path, achieving functional asynchronous visuomotor control on the low-cost Raspberry Pi 4. This represents a novel deployment strategy not demonstrated by prior GPU-centric VLA frameworks such as SmolVLA by Shukor et al. [22], whose work focused primarily on static robotic arms. Beyond proving technical feasibility, this work establishes a scalable methodology for deploying generalist robot intelligence under strict computational budgets.\n[...]\nParameter-efficient adaptation. We fine-tune a compact SmolVLM backbone using LoRA (rank 8, $\\alpha{=}8$ , dropout 0.1) to specialize visuomotor translation under tight memory/compute budgets (Alg. 1; Sec. III, pp. 2–3).\n[...]\n. 4).\n[...]\nLarge-scale multimodal systems such as PaLM-E, SayCan, and RT-2 have shown that unified language-conditioned reasoning enables robots to follow natural language commands and execute complex manipulation tasks. However, these approaches rely heavily on cloud-based computation and high-end GPUs, making them impractical for resource-limited or field-deployed robots. SMolVLA by Shukor et al. [22] introduced a small and efficient vision-language-action framework designed for community-driven robotic experimentation. It demonstrated that compact multimodal transformers could achieve competitive visuomotor reasoning performance while running on consumer-grade GPUs or CPUs. Nonetheles",
|
||||
"source_url": "https://arxiv.org/html/2511.05642",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2511.05642",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "A Vision-Language-Action Flow Model for General Robot Control",
|
||||
"snippet": "𝜋₀: A Vision-Language-Action Flow Model for General Robot Control\n[...]\n# $\\pi_{0}$ : A Vision-Language-Action Flow Model for General Robot Control\n[...]\nholds tremendous promise to unlock\n[...]\nfull potential of flexible, general, and dexterous\n[...]\nsystems, as well as to address some of\n[...]\ndeepest questions in artificial intelligence\n[...]\nHowever, bringing robot learning to\n[...]\nlevel of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how\n[...]\ncan design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of\n[...]\nability to perform tasks via direct prompting, follow language instructions from people and from a high-level VLM policy, and\n[...]\nability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.\n[...]\nIn this paper, we present a prototype model and learning framework, which we call $\\pi_",
|
||||
"source_url": "https://arxiv.org/abs/2410.24164",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2410.24164",
|
||||
"_exa_published_date": "2024-10-31T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "π: A Vision-Language-Action Flow Model for General Robot Control | OpenReview",
|
||||
"snippet": "π: A Vision-Language-Action Flow Model for General Robot Control | OpenReview\n[...]\n## π: A Vision-Language-Action Flow Model for General Robot Control\n[...]\nAbstract: Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.",
|
||||
"source_url": "https://openreview.net/forum?id=38a45ho9Nq",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=38a45ho9Nq",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "π0: A Vision-Language-Action Flow Model for General Robot Control",
|
||||
"snippet": "# π0: A Vision-Language-Action Flow Model for General Robot Control\n[...]\n```\n@article{Black20240AV,\n title={$\\pi$0: A Vision-Language-Action Flow Model for General Robot Control},\n author={Kevin Black and Noah Brown and Danny Driess and Adnan Esmail and Michael Equi and Chelsea Finn and Niccolo Fusai and Lachy Groom and Karol Hausman and Brian Ichter and Szymon Jakubczak and Tim Jones and Liyiming Ke and Sergey Levine and Adrian Li-Bell and Mohith Mothukuri and Suraj Nair and Karl Pertsch and Lucy Xiaoyang Shi and James Tanner and Quan Vuong and Anna Walling and Haohuan Wang and Ury Zhilinsky},\n journal={ArXiv},\n year={2024},\n volume={abs/2410.24164},\n url={https://api.semanticscholar.org/CorpusID:273811174}\n}\n\n```\n[...]\n- Kevin Black, Noah Brown, +21 authors Ury Zhilinsky\n[...]\n- Published in arXiv.org 31 October\n[...]\n2024\n[...]\n- Computer Science, Engineering\n[...]\nTLDR\n\nA novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge is proposed and evaluated in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning.Expand",
|
||||
"source_url": "https://www.semanticscholar.org/paper/%CF%800%3A-A-Vision-Language-Action-Flow-Model-for-General-Black-Brown/7e7e59d2e247d99954081080ddd5aae93d10b9e0",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://www.semanticscholar.org/paper/%CF%800%3A-A-Vision-Language-Action-Flow-Model-for-General-Black-Brown/7e7e59d2e247d99954081080ddd5aae93d10b9e0",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "𝜋₀: A Vision-Language-Action Flow Model for General Robot Control",
|
||||
"snippet": "𝜋₀: A Vision-Language-Action Flow Model for General Robot Control\n[...]\n# $\\pi_{0}$ : A Vision-Language-Action Flow Model for General Robot Control\n[...]\nRobot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence\n[...]\nHowever, bringing robot learning to\n[...]\nlevel of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and\n[...]\nability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.\n[...]\nIn this paper, we present a prototype model and learning framework, wh",
|
||||
"source_url": "https://arxiv.org/html/2410.24164v1",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2410.24164v1",
|
||||
"_exa_published_date": "2024-10-31T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "We train three policy classes: MolmoBot, a Molmo2-based multi-frame vision-language model with\n[...]\na flow-matching action head; MolmoBot-Pi0, which replicates the π0 architecture to enable direct\n[...]\ncomparison; and MolmoBot-SPOC, a lightweight policy suitable for edge deployment and amenable to\n[...]\nNVIDIA’s GR00T [1], Physical Intelligence’s π0 [2, 3], and Google DeepMind’s Gemini Robotics [4] frames\n[...]\nlarge-scale real-world training as the basis for generalist manipulation agents that act in the physical world.\n[...]\nMolmoBot-Pi0\n[...]\nUsing this data, we train three policy classes. Our flagship model, MolmoBot, is built on top of Molmo2 [7], our\n[...]\nvideo-language model capable of ingesting past frames for context. We augment this architecture with a DiT\u0002based flow-matching action head that is layerwise coupled to the vision-language backbone. Each action layer\n[...]\ncross-attends to the corresponding intermediate hidden states of the underlying VLM, while also incorporating\n[...]\nrobot-state features, allowing actions to be generated from multi-scale multimodal representations.\n[...]\nfrom MolmoBot, we also train MolmoBot-Pi0, which exactly replicates the π0 architecture for controlled\n[...]\ncomparison; and MolmoBot-SPOC, a lightweight non-VLA policy suitable for edge deployment and future RL\n[...]\nfine-tuning.\n[...]\nWe provide ablations demonstrating\n[...]\nimportance of data scale and diversity, and show through MolmoBot\u0002Pi0 that our\n[...]\nyields strong perfor",
|
||||
"source_url": "https://arxiv.org/pdf/2603.16861",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2603.16861",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Efficient Action Tokenization for Vision-Language-Action Models",
|
||||
"snippet": "Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, such models require us to choose a tokenization of our continuous action signals, which determines how the discrete symbols predicted by the model map to continuous robot actions. We find that current approaches for robot action tokenization, based on simple per-dimension, per-timestep binning schemes, typically perform poorly when learning dexterous skills from high-frequency robot data. To address this challenge, we propose a new compression-based tokenization scheme for robot actions, based on the discrete cosine transform. Our tokenization approach, Frequency-space Action Sequence Tokenization (FAST), enables us to train autoregressive VLAs for highly dexterous and high-frequency tasks where standard discretization methods fail completely. Based on FAST, we release FAST+, a universal robot action tokenizer, trained on 1M real robot action trajectories. It can be used as a black-box tokenizer for a wide range of robot action sequences, with diverse action spaces and control frequencies. Finally, we show that, when combined with the $\\bm{\\pi_{0}}$ VLA, our method can scale to training on 10k hours of robot data and match the performance of diffusion VLAs, while reducing training time by up to 5x.\n[...]\nFigure 4: Overview of the FAST action tokenization pipeline. Given a normalized c",
|
||||
"source_url": "https://arxiv.org/abs/2501.09747",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2501.09747",
|
||||
"_exa_published_date": "2025-01-16T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "FAST: Efficient Action Tokenization for Vision-Language ... - arXiv",
|
||||
"snippet": "Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors. However, such models require us to choose a tokenization of our continuous action signals, which determines how the discrete symbols predicted by the model map to continuous robot actions. We find that current approaches for robot action tokenization, based on simple per-dimension, per-timestep binning schemes, typically perform poorly when learning dexterous skills from high-frequency robot data. To address this challenge, we propose a new compression-based tokenization scheme for robot actions, based on the discrete cosine transform. Our tokenization approach, Frequency-space Action Sequence Tokenization (FAST), enables us to train autoregressive VLAs for highly dexterous and high-frequency tasks where standard discretization methods fail completely. Based on FAST, we release FAST+, a universal robot action tokenizer, trained on 1M real robot action trajectories. It can be used as a black-box tokenizer for a wide range of robot action sequences, with diverse action spaces and control frequencies. Finally, we show that, when combined with the $\\bm{\\pi_{0}}$ VLA, our method can scale to training on 10k hours of robot data and match the performance of diffusion VLAs, while reducing training time by up to 5x.\n[...]\nFigure 4: Overview of the FAST action tokenization pipeline. Given a normalized c",
|
||||
"source_url": "https://arxiv.org/html/2501.09747v1",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2501.09747v1",
|
||||
"_exa_published_date": "2025-01-16T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "ActionCodec: What Makes for Good Action Tokenizers",
|
||||
"snippet": "without any robotics\n[...]\nintroduce ActionCodec, a robust action tokenizer that integrates the\n[...]\n. Moreover, ActionCodec leverages Residual Vector Quantization (RVQ) (Lee et al., 2022) post-training to refine reconstruction fidelity and incorporates embodiment-specific soft prompts to facilitate knowledge transfer across diverse robotic platforms\n[...]\nthat VLA\n[...]\nActionCodec, without any additional architectural modifications\n[...]\nefficiency, success rates\n[...]\n. ActionCodec achieves SOTA performance in both\n[...]\nenvironments, providing a systematic\n[...]\nfor the future of VQ-\n[...]\nOur contributions are\n[...]\nas follows:\n[...]\nTokenization Schemes\n[...]\n20\n[...]\nsuffers from low training efficiency and ignores the\n[...]\nparallel decoding (\n[...]\n(Goy\n[...]\n., 2025),\n[...]\nfundamental inefficiencies of heuristic binning. Other\n[...]\nrepresent actions as strings for direct VLM prediction (Hancock et\n[...]\nhowever, this approach\n[...]\nno significant performance benefits while greatly increasing the token budget and extending latency to several seconds, limiting\n[...]\nPertsch et al., 2025) introduces Byte-Pair Encoding (BPE) on frequency-domain signals; however, its reliance on fixed geometric priors limits its capacity for cross-embodiment knowledge transfer. Data-driven schemes, particularly those based on Vector Quantization (VQ) (Wang et al., 2025b; Belkhale and Sadigh, 2024; Mete et al., 2024; Lee et al., 2024), offer a more flexible alternative by learning disc",
|
||||
"source_url": "https://arxiv.org/abs/2602.15397",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2602.15397",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "OAT: Ordered Action Tokenization",
|
||||
"snippet": "action tokenization\n[...]\nTo bridge this gap, we introduce Ordered Action Tokenization (OAT), a learned action tokenizer that discretizes continuous action chunks into highly compressed and causally ordered token sequences. OAT employs transformer-based register tokens to aggregate temporal information, finite scalar quantization (FSQ) to construct a discrete bottleneck, and nested dropout to explicitly induce ordering that aligns the latent space with autoregressive generation. The resulting tokenization ensures that any token prefix corresponds to a plausible action chunk. Beyond improved modelability, the ordered structure learned by OAT enables a key capability absent from prior approaches: prefix-based decoding. Autoregressive policies may terminate generation early and still produce valid actions, yielding a natural trade-off between computation and action fidelity. As additional tokens are generated, decoded actions are progressively refined.\n[...]\nAn alternative line of work explores frequency-domain compression, for instance Frequency-space Action Sequence Tokenization (FAST) [49], which employs the Discrete Cosine Transform (DCT) to decompose action chunks into frequency coefficients, followed by Byte Pair Encoding (BPE) [18]. FAST achieves high information density (P.1), and crucially, its low-frequency components first then high-frequency components ordering (P.3) improves downstream autoregressive policies: early token predictions capture the overall trajectory s",
|
||||
"source_url": "https://arxiv.org/html/2602.04215",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2602.04215",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding",
|
||||
"snippet": "the above challenges, we present a novel parallel decoding framework for the mainstream VLA model with action chunking, called Parallel\n[...]\nfor VLA (PD-VLA). Fig. 1 illustrates the\n[...]\nconcept of our parallel decoding approach. Our\n[...]\naction decoding as a system of\n[...]\nsolved through parallel fixed-point iteration methods, e.g., Jacobi fix-point iteration method [32]. This approach preserves\n[...]\nimproving decoding speed\n[...]\nthat we only accelerate the decoding process\n[...]\nVLA inference.\n[...]\n, our method enables friendly\n[...]\ntraining-free acceleration without redesign and modification of models (\n[...]\n, our method\n[...]\nsynergy with existing acceleration\n[...]\nVarious acceleration strategies, including quantization [21] and token pruning [5], have been effectively applied to LLMs, yet they often fail to meet the stringent real-time requirements of action generation. Efforts to enhance efficiency have led to architectural modifications in VLA models, such as DeeR-VLA [43], which dynamically adjusts inference depth, and QAIL [33], which integrates quantization-aware training. Further innovations, like RoboMamba [25] and TinyVLA [41], replace traditional attention mechanisms or focus on developing lightweight models from the ground up, frequently necessitating model re-training and additional data collection. Meanwhile, VLA-Cache [42] selectively caches static tokens and recomputes only dynamic or task-relevant ones. FAST [34] proposes a compression-based toke",
|
||||
"source_url": "https://arxiv.org/html/2503.02310v2",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2503.02310v2",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware | OpenReview",
|
||||
"snippet": "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware | OpenReview\n[...]\n## Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware\n[...]\nAbstract: Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations.",
|
||||
"source_url": "https://openreview.net/forum?id=e8Eu1lqLaf",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=e8Eu1lqLaf",
|
||||
"_exa_published_date": "2023-07-09T06:52:40.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Learning Fine-Grained Bimanual Manipulation with Low-Cost ... - arXiv",
|
||||
"snippet": "Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations. Project website: tonyzhaozh.github.io/aloha\n[...]\nImitation learning algorithm. Tasks that require precision and visual feedback present a significant challenge for imitation learning, even with high-quality demonstrations. Small errors in the predicted action can incur large differences in the state, exacerbating the “co",
|
||||
"source_url": "https://arxiv.org/abs/2304.13705",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2304.13705",
|
||||
"_exa_published_date": "2023-04-23T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with Transformers",
|
||||
"snippet": "coordinated biman\n[...]\n. To address the\n[...]\narms, particularly for synchronized actions. Therefore, we propose a novel imitation learning architecture that predicts cooperative actions. We differentiate the architecture for both arms and add an intermediate encoder layer, Inter-Arm Coordinated transformer Encoder (IACE),\n[...]\nInter-Arm Coordinated transformer Encoder (IACE), that can adjust the synchronization and timing of potential bimanual movements against the encoders corresponding to each arm. Our overall model\n[...]\na local Transformer encoder for each\n[...]\narm trajectory, the IACE to facilitate learning biman\n[...]\nactions, and a Transformer decoder to\n[...]\nthe action chunk. We compare two types of Transformer decoders: split decoders and single decoders.\n[...]\nWe build our proposed models on the ACT model to design different encoder and decoder structures. In particular, we propose a new design called the inter-arm coordinated transformer Encoder (IACE), which helps synchronize and time the movements of both arms.\n[...]\npropose basic architectures that consist of encoders\n[...]\narm, designed to leverage the biman\n[...]\nfeatures the IACE, allowing the individual robot arms to learn their trajectories while simultaneously considering the state\n[...]\nThe model should focus on the corresponding wrist camera and joint values to determine the appropriate trajectory for each robot arm. Each arm is supported by its local encoder. Global information is also integrated t",
|
||||
"source_url": "https://arxiv.org/html/2503.13916",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2503.13916",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "ALPHA-𝛼 and Bi-ACT Are All You Need: Importance of Position and Force Information/Control for Imitation Learning of Unimanual and Bimanual Robotic Manipulation with Low-Cost System",
|
||||
"snippet": "Autonomous manipulation in everyday tasks requires flexible action generation to handle complex, diverse real-world environments, such as objects with varying hardness and softness. Imitation Learning (IL) enables robots to learn complex tasks from expert demonstrations. However, a lot of existing methods rely on position/unilateral control, leaving challenges in tasks that require force information/control, like carefully grasping fragile or varying-hardness objects. As the need for diverse controls increases, there are demand for low-cost bimanual robots that consider various motor inputs. To address these challenges, we introduce Bilateral Control-Based Imitation Learning via Action Chunking with Transformers(Bi-ACT) and”A” ”L”ow-cost ”P”hysical ”Ha”rdware Considering Diverse Motor Control Modes for Research in Everyday Bimanual Robotic Manipulation (ALPHA- $\\alpha$ ). Bi-ACT leverages bilateral control to utilize both position and force information, enhancing the robot’s adaptability to object characteristics such as hardness, shape, and weight. The concept of ALPHA- $\\alpha$ is affordability, ease of use, repairability, ease of assembly, and diverse control modes (position, velocity, torque), allowing researchers/developers to freely build control systems using ALPHA- $\\alpha$ . In our experiments, we conducted a detailed analysis of Bi-ACT in unimanual manipulation tasks, confirming its superior performance and adaptability compared to Bi-ACT without force control. Base",
|
||||
"source_url": "https://arxiv.org/html/2411.09942",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2411.09942",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Learning Fine-Grained Bimanual Manipulation with Low-Cost ...",
|
||||
"snippet": "�W7/��A�/�\u0003c�s^���u�\"N�f�|5\n[...]\nǨ,��~.��a(Ц�\u0007\u001e�Pb��%��\u0019��\u001a?�3���\u001eS��~����l���\u0012�j\u0012tf;m���3\n[...]\n��7��h�\n[...]\n4\u001e\u0013L\n[...]\n\u000f9\\����� \u0001\u001a�� ��'��T�G�15�f����-ZB3���4������\u0000�[��\u0006����]�ӦbQi�uK\u0003��g4� h⠯� �8�@S�B��\u0002˔.�X��\u0007��IJ)��H\u0007I�����A��\u0011W�d������\u001a��\n[...]\n�Z\u0002O�[�\u0018^j�M\u000f\u001d��wE��df�\u0014h��C\u0004�\\���xͫ\u0010SB�\b��7�\u001aNb�ܫ#~8�7MD\n[...]\n�\u0001�\n[...]\nʬ`��y�1�1\u001fBf���R\u001fT$,Υ;lv\u0015@+%�\u0000�O�;k�S�6\"�\u0002#�J\n[...]\n�B2�+�\u0017.�\u0004�i\u0015�C �\u0005 ��\u0003�a�G�T�\u001b�W�\u0015\u000eג��RM�;�\u0004?�9$\u0014&�\u0013zT��Ʌ��z=7-�ip��1S3�� ���+\u001fI\u0013��q^��\u0018:���.�3,�\"t�v�\u001c��(c2�M�f�~�\u0010'\u001c�ӝ�ӨBf�d&\u0004�\u001d���\u00042�\b� C\u001e\u0006SΩ�T�A�8M!Ulŏ�f-�\u0003\u0004ݩF�30�nR�6ϛ�]E\b�\u0003]\\r�\\�n�5�� Q�uϬZZ|\u0003S�D<�_(�DU�<�8�0�x�)\u0006p��h�T\u0016�L\u0006��ו�X��#�n^+^q�8e0\n[...]\n�U*c��U�OB\u0011:\u0013 *��. �D:\u001b��iځ\u0015 l:��&:�Y��bxހl̰���?���\u00149�M �^|�ʛ1H���6����|��p��37I\u0002\u0001Z棍�\u0010�B���!���\\�A\u0012K��d�r�T5�>�ƨ^�t\u001e���\\|�m\u0011�>poŅ���\u0012��#5�%w[�|+\u0016iX���fy���X�ޕ�ݔs��M�\u0014��:�9���1b�����7�\u0015=U\n[...]\n\u0012�m\u001d�����!� I4\n[...]\n��A<����䬆��ᯟ \u0006�\u0018&�x⥌�;���H� �,#n�\u0013��U?���x�ؔ��d�\u0001��]�8�HoJ�}�\u0007 �ܓ �=��Z�\u0005G��T?�\u0004㘨�\u0007\u001a�!��*ǻ�5��[��x�D��i\u001c�Gc��\\\u000f�%��2��c�\u001b�d��6=d��Sm�;��\u0004��ݏF\u001e�\b1T�\u0006 �\u0012ch\u0007��� \u000f�\bV^��ԥ�e��_P�,� \u0010h_�O����\u0012ٱj\u0004��J\u001f?\u0012\u0012y\u0017 :{Zs�\u001b�x\u001d \\_�W�!\u0006���Ь��p���U!���\u001d�lm�n�B��E�\u0018g��\u0012\u0014w�ɺ��8�\u0010��N�\u0011v�M����`�\u001dYD|��=ZI8E��\u001c~\\�Ϫ��1���K���(\u0006�\u0005�����&\u0015\u0010R�\u0012_\u001e��\u0000��F�{ ��\u001bm�&��u\\�\u0005��ݶ�Ng��\u0011��4!��\u0005�Kh��.��a�3o�.'���2\\��wi3� >)iJ��sR� 3 �FB��O��\u0018�7��\u0019劍o s,���*��P\u0012\u001e\u0018\"���`�G�\u000f�2��l$q��Y��O?r\u001e_�P{\u0001�ܷG\b ����`�6\u001f��z��.��+���m��t%9��#E!$JtJe�$8C6�ͥ�\u001e�Ù�OV�]��bnvi)u0�fHy'���\u0010a����_,�D�M&uǒ�|�s]�D%�E>\"\"6� ��$ �� A����C��Q\u0003�`֩������}��\u0011� ��K/b�D��:�З�G<�\u001cy+wj�|U�\u0000r�\u0016Le�5��G��N|\u0010�K���� �Yv`\u0000K� �J0",
|
||||
"source_url": "https://openreview.net/pdf/4abc35d9793e56c5b73634eaf903e2495311fbcf.pdf",
|
||||
"discovered_for": [
|
||||
"rw.vla"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/pdf/4abc35d9793e56c5b73634eaf903e2495311fbcf.pdf",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[1512.03385] Deep Residual Learning for Image Recognition - arXiv",
|
||||
"snippet": "[1512.03385] Deep Residual Learning for Image Recognition\n[...]\n# Deep Residual Learning for Image Recognition\n[...]\nKaiming He Xiangyu Zhang Shaoqing Ren Jian Sun Microsoft Research {kahe, v-xiangz, v-shren, jiansun}@microsoft.com\n[...]\nDeeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers—8 $\\times$ deeper than VGG nets [41] but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers.\n[...]\nreferenced mapping. To\n[...]\nextreme, if\n[...]\nwould be easier\n[...]\nresidual to zero than\n[...]\nnonlinear layers.\n[...]\non ImageNet [36\n[...]\nOn the ImageNet classification dataset [36], we obtain excellent results by extremely deep residual nets. Our 152-layer residual net is the deepest network ever presented on ImageNet, while still having lower complexity than VGG nets [41]. Our ensem",
|
||||
"source_url": "https://arxiv.org/abs/1512.03385",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1512.03385",
|
||||
"_exa_published_date": "2015-12-10T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[1512.03385] Deep Residual Learning for Image Recognition",
|
||||
"snippet": "[1512.03385] Deep Residual Learning for Image Recognition\n[...]\n# Deep Residual Learning for Image Recognition\n[...]\nKaiming He Xiangyu Zhang Shaoqing Ren Jian Sun Microsoft Research {kahe, v-xiangz, v-shren, jiansun}@microsoft.com\n[...]\nDeeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers—8 $\\times$ deeper than VGG nets [41] but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers.\n[...]\nreferenced mapping. To\n[...]\nextreme, if\n[...]\nwould be easier\n[...]\nresidual to zero than\n[...]\nnonlinear layers.\n[...]\non ImageNet [36\n[...]\nOn the ImageNet classification dataset [36], we obtain excellent results by extremely deep residual nets. Our 152-layer residual net is the deepest network ever presented on ImageNet, while still having lower complexity than VGG nets [41]. Our ensem",
|
||||
"source_url": "https://arxiv.org/abs/1512.03385v1",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1512.03385v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Deep Residual Learning for Image Recognition - People | MIT CSAIL",
|
||||
"snippet": "Deep Residual Learning\nfor Image Recognition\nKaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun\nwork done at\nMicrosoft Research Asia\n[...]\nResNet @ ILSVRC & COCO 2015 Competitions\n[...]\n1st places in all five main tracks\n[...]\n• ImageNet Classification: “Ultra-deep” 152-layer nets \n• ImageNet Detection: 16% better than 2nd\n[...]\n• ImageNet Localization: 27% better than 2nd\n[...]\n• COCO Detection: 11% better than 2nd\n[...]\n• COCO Segmentation: 12% better than 2nd\n[...]\n*improvements are relative numbers\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nRevolution of Depth\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016.\n[...]\nKaiming He, Xiang",
|
||||
"source_url": "https://pdfs.semanticscholar.org/1cea/9b1931b9e87641708fec43d03f2a58f4d2b0.pdf",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://pdfs.semanticscholar.org/1cea/9b1931b9e87641708fec43d03f2a58f4d2b0.pdf",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Deep Residual Learning for Image Recognition: A Survey - MDPI",
|
||||
"snippet": "Deep Residual Learning for Image Recognition: A Survey\n[...]\n# Deep Residual Learning for Image Recognition: A Survey\n[...]\nMuhammad Shafiq\n[...]\n1,* and\n\nZhaoquan Gu\n[...]\n2,3,*\n[...]\nCyberspace Institute of Advanced Technology, Guangzhou University, Guangzhou 510006, China\n[...]\nDepartment of New Networks, Peng Cheng Laboratory, Shenzhen 518055, China\n[...]\nDepartment of Computer Science and Technology, Harbin Institute of Technology, Shenzhen 518055, China\n[...]\nAppl. Sci. 2022, 12(18), 8972; https://doi.org/10.3390/app12188972\n[...]\nDeep Residual Networks have recently been shown to significantly improve the performance of neural networks trained on ImageNet, with results beating all previous methods on this dataset by large margins in the image classification task. However, the meaning of these impressive numbers and their implications for future research are not fully understood yet. In this survey, we will try to explain what Deep Residual Networks are, how they achieve their excellent results, and why their successful implementation in practice represents a significant advance over existing techniques. We also discuss some open questions related to residual learning as well as possible applications of Deep Residual Networks beyond ImageNet. Finally, we discuss some issues that still need to be resolved before deep residual learning can be applied on more complex problems.\n[...]\ndeep residual learning for image recognition\n[...]\ndeep residual learning;\n[...]\nDeep resid",
|
||||
"source_url": "https://www.mdpi.com/2076-3417/12/18/8972",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://www.mdpi.com/2076-3417/12/18/8972",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[1603.05027] Identity Mappings in Deep Residual Networks - arXiv",
|
||||
"snippet": "Kaiming He Xiangyu Zhang Shaoqing Ren Jian Sun\n[...]\nDeep residual networks [1] have emerged as a family of extremely deep architectures showing compelling accuracy and nice convergence behaviors. In this paper, we analyze the propagation formulations behind the residual building blocks, which suggest that the forward and backward signals can be directly propagated from one block to any other block, when using identity mappings as the skip connections and after-addition activation. A series of ablation experiments support the importance of these identity mappings. This motivates us to propose a new residual unit, which makes training easier and improves generalization. We report improved results using a 1001-layer ResNet on CIFAR-10 (4.62% error) and CIFAR-100, and a 200-layer ResNet on Image\n[...]\n. Code is available at: https://github.com/KaimingHe/resnet-1k-layers.\n[...]\n- [1] He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. (2016)",
|
||||
"source_url": "https://arxiv.org/abs/1603.05027",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1603.05027",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "(AttnRes) [30] recently showed that the fixed residual connections in Trans\u0002formers, which accumulate layer outputs with uniform unit weights, can be re\u0002placed by learned softmax attention over all preceding layer outputs. A single\n[...]\npseudo-query vector per layer selects which earlier representations to aggregate,\n[...]\nenabling content-aware, position-specific routing with minimal overhead. This\n[...]\nprinciple of learned\n[...]\nAttention Residuals (AttnRes) were proposed by Chen et al. [30] for Transformer\u0002based LLMs. The standard residual connection accumulates layer outputs with\n[...]\nfixed unit weights: hl = hl−1 + fl(hl−1). As network depth grows, this causes\n[...]\ntwo problems: feature dilution, where each layer’s relative contribution to the\n[...]\naccumulated sum diminishes, and unbounded magnitude growth, a well-known\n[...]\nissue in PreNorm Transformers. AttnRes replaces the fixed accumulation with\n[...]\nsoftmax attention over all preceding layer outputs, parameterized by a single\n[...]\nlearned pseudo-query wl ∈ R\n[...]\nd per layer. This enables selective, content-aware\n[...]\naccess to earlier representations while keeping output magnitudes bounded.\n[...]\nWe extend Attention Residuals from same-dimensional layer aggregation in\n[...]\nIn standard Transformers [31], the residual connection at layer l accumulates\n[...]\nwhere fl denotes the layer computation (e.g., self-attention or feed-forward net\u0002work). Attention Residuals [30] replace this fixed accumulation with a",
|
||||
"source_url": "https://arxiv.org/pdf/2604.03297",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2604.03297",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2603.15031] Attention Residuals - arXiv",
|
||||
"snippet": "Untitled Document\n\n$0$ $5$ $10$ $15$ $20$\n\nUntitled Document\n$0$ $5$ $10$ $15$ $20$\nBETA",
|
||||
"source_url": "https://arxiv.org/abs/2603.15031",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2603.15031",
|
||||
"_exa_published_date": "2026-03-16T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm",
|
||||
"snippet": "In this paper, we propose SiameseNorm, an elegant two\n[...]\nstream residual architecture that unifies\n[...]\nmaintain two residual streams\n[...]\nshared parameters:\n[...]\nthe advantages of\n[...]\nnegligible computational overhead\n[...]\nboosts accuracy from\n[...]\n128\n[...]\n639.6\n[...]\nwhere the product denotes an ordered composition of Jacobians from layer N−1N-1 down to i+1i+1. Notably, the term 𝐈\\mathbf{I} corresponds to the skip connection, which preserves an explicit identity gradient path. This allows gradients to flow through the network without explicit attenuation, facilitating the training of large scale models. However, it implicitly allows the representation magnitudes to grow unbounded. As noted previously, Pre-Norm exhibits insufficient effective depth, an issue that likely stems from a structural mismatch: As shown in Figure˜2(a), the main path accumulates residual updates without re-normalization, causing hidden state magnitudes to grow with depth (peri-ln). Consequently, deeper blocks encounter a scaling imbalance: they must influence an increasingly high-magnitude main path while being restricted to normalized, fixed-scale inputs. This growing disparity effectively dilutes the relative contribution of deeper layers, thereby limiting the effective depth of the model.\n[...]\nBy maintaining a clean identity path, Pre-Norm ensures stable gradient propagation. However, this comes at the cost of unbounded magnitude growth. As illustrated in Figure 2(a), while the input ",
|
||||
"source_url": "https://arxiv.org/html/2602.08064v1",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/html/2602.08064v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Attention Residuals - arXiv",
|
||||
"snippet": "Untitled Document\n\n$0$ $5$ $10$ $15$ $20$\n\nUntitled Document\n$0$ $5$ $10$ $15$ $20$\nBETA",
|
||||
"source_url": "https://arxiv.org/pdf/2603.15031",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2603.15031",
|
||||
"_exa_published_date": "2026-03-16T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "However, deep PreNorm models exhibit their own limita\u0002tions. The clean residual path can lead to representation\n[...]\nlayers fail to learn new features (Li\n[...]\n). This “Curse of Depth” is characterized by\n[...]\nexponential growth in activation variance and layers that de\u0002volve into identity functions, limiting the benefits of scaling\n[...]\nPre-LayerNorm (PreNorm): To address the training in\u0002stability of PostNorm, especially when training deep models,\n[...]\n. Here,\n[...]\neach sub-layer,\n[...]\n. This creates a clean, identity\n[...]\nover PreNorm\n[...]\nRepresentation Collapse A primary pathology in deep\n[...]\nPreNorm Transformers is the uncontrolled growth of fea\u0002ture variance, which leads to representation collapse. In a\n[...]\nstandard PreNorm block, the main residual path acts as an\n[...]\nidentity map, causing variance to accumulate linearly with\n[...]\ndepth. Formally, for a network of depth L, the variance of\n[...]\nthe hidden states scales as Var(X′\n[...]\nΘ(L) (Kedia et al.,\n[...]\n024).\n[...]\nThis variance explosion degrades the learning capability of\n[...]\ndeep layers. Consider the Jacobian of the l-th PreNorm\n[...]\nblock. Let Res(·) denote the transformation within the resid\u0002ual branch (e.g., Self-Attention or FFN). Since the transfor\u0002mation within the residual branch operates on inputs nor\u0002malized by the feature standard deviation σl =\n[...]\nthe Jacobian JPre can be expressed as:\n[...]\nAs l → ∞, σl → ∞, causing the residual term to vanish at a\n[...]\nrate of O(1/\n[...]\nl).",
|
||||
"source_url": "https://arxiv.org/pdf/2601.22580v1",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/2601.22580v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[2409.19606] Hyper-Connections - arXiv",
|
||||
"snippet": "We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed in residual connection variants, such as the seesaw effect between gradient vanishing and representation collapse. Theoretically, hyper-connections allow the network to adjust the strength of connections between features at different depths and dynamically rearrange layers. We conduct experiments focusing on the pre-training of large language models, including dense and sparse models, where hyper-connections show significant performance improvements over residual connections. Additional experiments conducted on vision tasks also demonstrate similar improvements. We anticipate that this method will be broadly applicable and beneficial across a wide range of AI problems.\n[...]\nDeep learning has achieved tremendous success across various domains, where residual connections (He et al., 2016) have been instrumental in contemporary neural network architectures, including transformers and CNNs. Residual connections help mitigate the problem of gradient vanishing, enabling the effective training of very deep networks. However, it is important to acknowledge that residual connections are not infallible solutions and still present limitations that remain unresolved.\n[...]\nDriven by the limitations of residual connections, an important question arises: Can neural networks autonomously learn the optimal streng",
|
||||
"source_url": "https://arxiv.org/abs/2409.19606",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/2409.19606",
|
||||
"_exa_published_date": "2024-09-29T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Hyper-Connections - OpenReview",
|
||||
"snippet": "Hyper-Connections | OpenReview\n\n## Hyper-Connections\n\n### Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou\n\nICLR 2025 Postereveryonesince 04 Oct 2024\">Everyone Revisions BibTeX CC BY 4.0\n\nKeywords: Network Architecture, Residual Connections, LLMs, Pre-training\n\nAbstract: We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed in residual connection variants, such as the seesaw effect between gradient vanishing and representation collapse. Theoretically, hyper-connections allow the network to adjust the strength of connections between features at different depths and dynamically rearrange layers. We conduct experiments focusing on the pre-training of large language models, including dense and sparse models, where hyper-connections show significant performance improvements over residual connections. Additional experiments conducted on vision tasks also demonstrate similar improvements. We anticipate that this method will be broadly applicable and beneficial across a wide range of AI problems.\n\nPrimary Area: foundation or frontier models, including LLMs\n\nCode Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics.\n\nSubmission Guidelines: I certify that this submission complies with the submission instructions as described on https://iclr.cc/Con",
|
||||
"source_url": "https://openreview.net/forum?id=9FqARW7dwB",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=9FqARW7dwB",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Birkhoff-Exact Hyper-Connections: Exact Spectral Stability for Deep Residual Networks | OpenReview",
|
||||
"snippet": "## Birkhoff-Exact Hyper-Connections: Exact Spectral Stability for Deep Residual Networks\n[...]\nKeywords: doubly stochastic matrices, spectral stability, deep residual networks, Birkhoff-von Neumann theorem, hyper-connections, token mixing, extreme depth training, quantization robustness\n[...]\nTL;DR: We propose BE-HC, which uses the Birkhoff-von Neumann theorem to construct exactly doubly stochastic mixing matrices as convex combinations of permutation matrices, enabling stable training at 1000+ layers where prior methods fail.\n[...]\nAbstract: Learnable information routing in deep networks faces the *depth-stability-efficiency trilemma*: architectures that scale to extreme depths often sacrifice efficiency; efficient approaches lack stability guarantees. Prior work uses iterative Sinkhorn-Knopp normalization to approximate doubly stochastic mixing matrices, but residual errors destabilize training beyond several hundred layers. We propose **Birkhoff-Exact Hyper-Connections (BE-HC)**, which leverages the Birkhoff-von Neumann theorem to construct *exactly* doubly stochastic matrices as convex combinations of permutation matrices. This guarantees spectral radius $\\rho = 1$ exactly—not approximately—enabling stable training at unprecedented depths. **Key results:** (1) *Extreme depth:* BE-HC trains stably at **1000 layers**, achieving 35.71% accuracy where ReZero and other baselines fail to converge. (2) *Long context:* BE-HC handles **8K tokens** on a single V100 GPU (22.56% vali",
|
||||
"source_url": "https://openreview.net/forum?id=jpIjkN1B1Q",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=jpIjkN1B1Q",
|
||||
"_exa_published_date": "2026-03-02T22:01:32.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "We present the first study of Hyper-Connections (HC) for volumetric multi-modal\n[...]\nbrain tumor segmentation, integrating them as a drop-in replacement for fixed\n[...]\nresidual connections across five architectures: nnU-Net, SwinUNETR, VT-UNet,\n[...]\nNetpp.\n[...]\nIn this work, we explore Hyper-Connections (HC) [18], a recently proposed generalization of residual\n[...]\nconnections that enables dynamic, input-dependent feature aggregation. While HC has shown\n[...]\nit has not been investigated for medical image segmentation, particularly in volumetric and multi\u0002modal settings. We extend HC\n[...]\nboth 3D\n[...]\nand training dynamics. Residual connections [5]\n[...]\nstable optimization of deep networks through\n[...]\narchitectural components. In contrast, the Hyper-Connection framework [18] provides a unified\n[...]\nformulation that simultaneously adapts both depth-wise and width-wise connections, representing a\n[...]\nstrict generalization of prior approaches. To the best of our knowledge, this work presents the first\n[...]\nHyper-Connections (HC) were originally proposed in the context of large language model pre-training,\n[...]\nwhere the primary challenge is to maintain stable gradient flow across very deep transformer archi\u0002tectures operating on sequential token embeddings [18]. In that setting, HC improves optimization\n[...]\nby learning adaptive depth-wise aggregation, thereby avoiding limitations associated with fixed\n[...]\nresidual formulations such as Pre-Norm and Post-Norm va",
|
||||
"source_url": "https://www.arxiv.org/pdf/2603.19844",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://www.arxiv.org/pdf/2603.19844",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections",
|
||||
"snippet": "Multi-stream transformer architectures have recently been proposed as a promising direction for managing representation collapse and the vanishing gradient problem for residual connections, yet their internal mechanisms remain unexplored. In particular, the recently introduced Manifold-Constrained Hyper-Connections (mHC) architecture posits multiple residual streams with constrained interaction, but lacks in-depth mechanistic analysis. We present the first open-source mHC language model (https://huggingface.co/wgpeng/mhc-780m) and analyze the multiple-stream architecture with a suite of representation-level metrics and causal interventions to probe how parallel streams encode and utilize information. Specifically, we introduce a systematic stream ablation-and-rescue framework that enables direct causal comparison of residual streams during inference. Through targeted pairwise interventions and controlled recovery experiments, we distinguish functional redundancy from asymmetric utilization and reveal how information is distributed across streams beyond what is observable from representational similarity alone.\n[...]\nHyper-Connections extend the standard transformer residual architecture by allowing multiple residual streams per layer, dynamically mixed through learned routing matrices (He et al., 2015; Zhu et al., 2025). Manifold-Constrained Hyper-Connections (mHC) further refines this framework by imposing geometric constraints on inter-stream mixing (Xie et al., 2026).\n[...",
|
||||
"source_url": "https://www.arxiv.org/pdf/2603.14833",
|
||||
"discovered_for": [
|
||||
"rw.attnres"
|
||||
],
|
||||
"_exa_id": "https://www.arxiv.org/pdf/2603.14833",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "DECOUPLED WEIGHT DECAY REGULARIZATION\n[...]\nIlya Loshchilov & Frank Hutter\n[...]\nL2 regularization and weight decay regularization are equivalent for standard\n[...]\nstochastic gradient descent (when rescaled by the learning rate), but as we demon\u0002strate this is not the case for adaptive gradient algorithms, such as Adam. While\n[...]\nexpose), we propose a simple modification to recover the original formulation of\n[...]\nweight decay regularization by decoupling the weight decay from the optimization\n[...]\nsteps taken w.r.t. the loss function. We provide empirical evidence that our pro\u0002posed modification (i) decouples the optimal choice of weight decay factor from\n[...]\nthe setting of the learning rate for both standard SGD and Adam and (ii) substan\u0002tially improves Adam’s generalization performance, allowing it to compete with\n[...]\nSGD with momentum on image classification datasets (on which it was previously\n[...]\ntypically outperformed by the latter). Our proposed decoupled weight decay has\n[...]\ncommunity has implemented\n[...]\nit in TensorFlow and PyTorch; the complete source code for our experiments is\n[...]\navailable at https://github.com/loshchil/AdamW-and-SGDW\n[...]\nThe main contribution of this paper is to improve regularization in Adam by decoupling the weight\n[...]\ndecay from the gradient-based update. In a comprehensive analysis, we show that Adam generalizes\n[...]\nsubstantially better with decoupled weight decay than with L2 regularization, achieving 15% relative\n[.",
|
||||
"source_url": "https://arxiv.org/pdf/1711.05101v3",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/1711.05101v3",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[1711.05101] Decoupled Weight Decay Regularization - arXiv",
|
||||
"snippet": "[1711.05\n[...]\npled Weight Decay Regularization\n[...]\nIlya Loshchilov & Frank Hutter University of Freiburg Freiburg, Germany, {ilya,fh}@cs.uni-freiburg.de\n[...]\nL2 regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is not the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L2 regularization (often calling it “weight decay” in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by decoupling the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam’s generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at https://github.com/loshchil/AdamW-and-SGDW\n[...]\nThe main contribution of this paper is to improve regularization in Adam by decoupling ",
|
||||
"source_url": "https://arxiv.org/abs/1711.05101",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1711.05101",
|
||||
"_exa_published_date": "2017-11-14T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "[PDF] Decoupled Weight Decay Regularization - arXiv",
|
||||
"snippet": "DECOUPLED WEIGHT DECAY REGULARIZATION\n[...]\nIlya Loshchilov & Frank Hutter\n[...]\nL2 regularization and weight decay regularization are equivalent for standard\n[...]\nstochastic gradient descent (when rescaled by the learning rate), but as we demon\u0002strate this is not the case for adaptive gradient algorithms, such as Adam. While\n[...]\nexpose), we propose a simple modification to recover the original formulation of\n[...]\nweight decay regularization by decoupling the weight decay from the optimization\n[...]\nsteps taken w.r.t. the loss function. We provide empirical evidence that our pro\u0002posed modification (i) decouples the optimal choice of weight decay factor from\n[...]\nthe setting of the learning rate for both standard SGD and Adam and (ii) substan\u0002tially improves Adam’s generalization performance, allowing it to compete with\n[...]\nSGD with momentum on image classification datasets (on which it was previously\n[...]\ntypically outperformed by the latter). Our proposed decoupled weight decay has\n[...]\ncommunity has implemented\n[...]\nit in TensorFlow and PyTorch; the complete source code for our experiments is\n[...]\navailable at https://github.com/loshchil/AdamW-and-SGDW\n[...]\nThe main contribution of this paper is to improve regularization in Adam by decoupling the weight\n[...]\ndecay from the gradient-based update. In a comprehensive analysis, we show that Adam generalizes\n[...]\nsubstantially better with decoupled weight decay than with L2 regularization, achieving 15% relative\n[.",
|
||||
"source_url": "https://arxiv.org/pdf/1711.05101",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/pdf/1711.05101",
|
||||
"_exa_published_date": "2019-01-04T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "Decoupled Weight Decay Regularization - OpenReview",
|
||||
"snippet": "Decoupled Weight Decay Regularization | OpenReview\n\n## Decoupled Weight Decay Regularization\n\nICLR 2019 Conference Blind SubmissionReaders: Everyone\n\nAbstract: L$_2$ regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \\emph{not} the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L$_2$ regularization (often calling it ``weight decay'' in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by \\emph{decoupling} the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam's generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at \\url{https://github.com/loshchil/AdamW-and-SGDW}\n\nKeywords: optimization, regularization, weight decay, Adam\n\nCode: ",
|
||||
"source_url": "https://openreview.net/forum?id=Bkg6RiCqY7",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/forum?id=Bkg6RiCqY7",
|
||||
"_exa_published_date": "2018-09-27T08:45:35.000Z"
|
||||
},
|
||||
{
|
||||
"title": "",
|
||||
"snippet": "Ilya Loshchilov & Frank Hutter \nUniversity of Freiburg \nFreiburg, Germany, \nilya.loshchilov@gmail.com, fh@cs.uni-freiburg.de \n\nABSTRACT\n[...]\nL2 regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demon strate this is not the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L2 regularization (often calling it “weight decay” in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by decoupling the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our pro posed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substan tially improves Adam’s generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at https://github.com/loshchil/AdamW-and-SGDW\n[...]\n(AdamW),\n[...]\n(Loshchilov & Hutter,\n[...]\n. Figure 1\n[...]\nweight decay (\n[...]\nshow that weight decay ",
|
||||
"source_url": "https://openreview.net/pdf?id=Bkg6RiCqY7",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://openreview.net/pdf?id=Bkg6RiCqY7",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "PyTorch: An Imperative Style, High-Performance Deep Learning Library",
|
||||
"snippet": "PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\nDeep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several commonly used benchmarks.",
|
||||
"source_url": "https://papers.nips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://papers.nips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[PDF] An Imperative Style, High-Performance Deep Learning Library - NIPS",
|
||||
"snippet": "PyTorch: An Imperative Style, High-Performance Deep Learning Library\n\n| | Adam | Paszke | | Sam | Gross | | | Francisco | Massa | | |\n[...]\nAbstract Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several common benchmarks.\n[...]\nWith the increased interest in deep learning in recent years, there has been an explosion of machine learning tools. Many popular frameworks such as Caffe [1], CNTK [2], TensorFlow [3], and Theano [4], construct a static dataflow graph that represents the computation and which can then be applied repeatedly to batches of data. This approach provides visibility into the whole comput",
|
||||
"source_url": "https://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "[1912.01703v1] PyTorch: An Imperative Style, High-Performance Deep Learning Library",
|
||||
"snippet": "[1912.01703v1] PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n# Title:PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n> Abstract:Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several common benchmarks.",
|
||||
"source_url": "https://arxiv.org/abs/1912.01703v1",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1912.01703v1",
|
||||
"_exa_published_date": null
|
||||
},
|
||||
{
|
||||
"title": "PyTorch: An Imperative Style, High-Performance Deep Learning ...",
|
||||
"snippet": "[1912.01703] PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\n# PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\nAdam Paszke University of Warsaw adam.paszke@gmail.com Sam Gross Facebook AI Research sgross@fb.com Francisco Massa Facebook AI Research fmassa@fb.com Adam Lerer Facebook AI Research alerer@fb.com James Bradbury Google jekbradbury@gmail.com Gregory Chanan Facebook AI Research gchanan@fb.com Trevor Killeen Self Employed killeent@cs.washington.edu Zeming Lin Facebook AI Research zlin@fb.com Natalia Gimelshein NVIDIA ngimelshein@nvidia.com Luca Antiga Orobix luca.antiga@orobix.com Alban Desmaison Oxford University alban@robots.ox.ac.uk Andreas Köpf Xamla andreas.koepf@xamla.com Edward Yang Facebook AI Research ezyang@fb.com Zach DeVito Facebook AI Research zdevito@cs.stanford.edu Martin Raison Nabla martinraison@gmail.com Alykhan Tejani Twitter atejani@twitter.com Sasank Chilamkurthy Qure.ai sasankchilamkurthy@gmail.com Benoit Steiner Facebook AI Research benoitsteiner@fb.com Lu Fang Facebook lufang@fb.com Junjie Bai Facebook jbai@fb.com Soumith Chintala Facebook AI Research soumith@gmail.com\n[...]\nDeep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popu",
|
||||
"source_url": "https://arxiv.org/abs/1912.01703",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "https://arxiv.org/abs/1912.01703",
|
||||
"_exa_published_date": "2019-12-03T00:00:00.000Z"
|
||||
},
|
||||
{
|
||||
"title": "PyTorch: An Imperative Style, High-Performance Deep Learning Library",
|
||||
"snippet": "PyTorch: An Imperative Style, High-Performance Deep Learning Library\n[...]\nDeep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several commonly used benchmarks.",
|
||||
"source_url": "http://papers.neurips.cc/paper/9015-pytorch-an-",
|
||||
"discovered_for": [
|
||||
"method.training"
|
||||
],
|
||||
"_exa_id": "http://papers.neurips.cc/paper/9015-pytorch-an-",
|
||||
"_exa_published_date": null
|
||||
}
|
||||
]
|
||||
}
|
||||
Reference in New Issue
Block a user