Glass crashes slashed? Ant Group embodied AI unit claims breakthrough in robot sensing
Ant Group's Robbyant unit has unveiled new AI vision models, LingBot-Depth 2.0 and LingBot-Vision, designed to improve robot perception of transparent objects and complex environments.
Intelligence analysis by Gemini 2.5 Flash

The Chinese fintech giant's embodied AI arm claims its new foundational visual model, LingBot-Vision, can accurately perceive glass, mirrors, and transparent objects, a long-standing challenge for robots. It reportedly outperforms Meta Platforms' DINOv3 on benchmarks, using significantly fewer parameters and less training data, signaling a potential leap in efficient robot navigation.
Imagine a robot trying to clean a glass table. Usually, it might bump into it because glass is hard to see. But Ant Group made a special 'eye' for robots that helps them see the edges of clear things, like drawing an invisible line around them. It's like giving the robot special glasses that make clear objects glow a little, so it knows where they are and doesn't crash.
Analysis
The Challenge of Transparent Perception
Robots have long struggled with accurately perceiving transparent and reflective objects like glass and mirrors. These materials often confuse traditional vision systems, leading to collisions and hindering a robot's ability to navigate and interact safely within complex, real-world environments. This limitation has been a significant bottleneck for the widespread deployment of embodied AI, where machines need to understand and operate in physical spaces with human-like dexterity and awareness.
Ant Group's Robbyant unit, also known as Ant Lingbo Technology, has directly addressed this issue with its new LingBot-Depth 2.0 and LingBot-Vision models. The firm claims these technologies are specifically trained to recognize object edges with high precision, down to a fraction of a pixel. This enhanced spatial understanding is crucial for robots to differentiate between empty space and transparent barriers, potentially preventing accidents and enabling more sophisticated tasks.
Efficiency and Performance Against Rivals
LingBot-Vision enters a competitive field, notably challenging Meta Platforms' open-source DINOv3 vision model. What sets Ant Group's claim apart is its focus on structural efficiency. According to a research paper by the Robbyant team, LingBot-Vision surpassed the 7-billion-parameter DINOv3 across multiple metrics on the NYUv2 depth-estimation benchmark. This was achieved using only one-seventh as many parameters and less than a third of the training data.
This efficiency is a critical factor in AI development, as it suggests that powerful perception capabilities can be achieved with fewer computational resources. Such advancements could make sophisticated embodied AI more accessible and scalable, reducing the energy and hardware requirements for deploying advanced robotic systems. The ability to achieve superior performance with less data and computational overhead represents a significant step forward in the pursuit of more practical and sustainable AI solutions for robotics.
Implications for Embodied AI and Robotics
The development of LingBot-Vision and LingBot-Depth 2.0 has profound implications for the future of embodied AI and robotics. By enabling robots to "accurately and stably" see in unpredictable environments, these models could unlock new applications across various industries. From logistics and manufacturing to service robots operating in homes and public spaces, improved perception of transparent objects means greater safety, reliability, and versatility for autonomous machines.
This breakthrough could accelerate the integration of robots into daily life, allowing them to perform tasks that were previously too hazardous or complex due to visual ambiguities. The ability to precisely map 3D spaces and identify subtle boundaries could lead to more nuanced interactions between robots and their surroundings, paving the way for more intelligent and adaptable robotic assistants. As AI labs worldwide race to equip machines with advanced cognitive abilities, Ant Group's contribution highlights the ongoing progress in making robots truly aware of their physical world.
Key points
- Ant Group's Robbyant unit launched new AI vision models, LingBot-Depth 2.0 and LingBot-Vision.
- The models aim to solve the long-standing challenge of robots accurately perceiving glass, mirrors, and transparent objects.
- LingBot-Vision reportedly outperforms Meta Platforms' DINOv3 on benchmarks, using significantly fewer parameters and less training data.
- The technology is designed to pinpoint object boundaries with high precision, down to a fraction of a pixel.
- This advancement could enable robots to navigate complex physical spaces more safely and effectively.
This breakthrough could lead to safer and more capable robots that can navigate complex human environments without crashing into transparent objects. It may also accelerate the development of more efficient AI models, reducing the computational resources needed for advanced robotics.
While promising, the transition from benchmark success to widespread, robust real-world deployment often presents unforeseen challenges, including varied lighting conditions and complex environments not fully captured in training data.



