NVIDIA’s Vera Rubin architecture reportedly demonstrates a 30x improvement in throughput per watt compared to its Blackwell predecessor, along with a 35x reduction in token costs for agentic AI workloads, according to Wccftech. The performance gains suggest significant efficiency improvements in power consumption and computational cost for AI inference applications.