Bankless · Monday, August 31, 2026
Shahin Farshchi explains that current robots utilize a traditional AI workflow involving sensing, perception, planning, and actuation. While new Vision-Language Models (VLMs) and Vision-Language Action Models (VLAs) offer greater generalizability by using language to interpret scenes and generate actions, he believes a hybrid approach combining these with traditional methods will likely be the most effective for future robotics.
“Okay. So that is the, I would say, common AI-based workflow that exists in robots in the field today.”
“Now, what we're seeing with VLAs and VLMs is some flavor of, of using language to interpret a scene and then generating language from that language should then take some kind of action.”
“And it's my expectation that it's going to be some kind of hybrid process of the two, if that makes sense.”