🧮 Today's top-end GPUs are designed around AI, and they keep getting faster at rough arithmetic. Weather forecasting and materials science often need long, precise numbers instead. Katsuhisa Ozaki, a professor at Shibaura Institute of Technology in Japan, and his collaborators found a way to stack many rough calculations into one precise answer. As of September 2026, that method ships inside NVIDIA's CUDA Toolkit 13.4.
AI and science want different kinds of numbers
Simulations of weather, earthquakes or molecules have long relied on double precision to keep results accurate. Usually written FP64, it stores each number in 64 bits. AI training and inference, by contrast, mostly get by with 8-bit formats such as FP8 and INT8. Over the past few years, GPU makers have spent more and more of each chip on making those low-precision formats fast.
The result is an awkward reversal. According to Shibaura Institute of Technology, NVIDIA's B200 delivers about 40 TFLOPS of double-precision performance (one TFLOPS is a trillion calculations per second). Its successor, Blackwell Ultra (B300), manages 1.2 to 1.3 TFLOPS, less than one-thirtieth of that.
Stacking rough math, one remainder at a time
The Ozaki scheme started with a 2012 paper. It splits large numbers into small pieces that low-precision hardware can multiply exactly, then adds the partial products back together, roughly the way long multiplication on paper works digit by digit. This first version, now called Ozaki Scheme I, went into CUDA 13.0 Update 2 in October 2025 as an opt-in feature that developers switch on.
Scheme I's weakness is cost. The finer you cut the numbers to gain precision, the more multiplications you need, growing roughly with the square of the precision.
Ozaki Scheme II, which Ozaki proposed in 2025 with Yuki Uchino and Toshiyuki Imamura of RIKEN, works on a different principle: the Chinese remainder theorem. For example, if you know a number's remainders when divided by 7, 11 and 13, and the number is below 1,001, there is only one possibility. Scheme II turns the matrix values into integers, has the AI-oriented units multiply only the remainders left over after dividing by a handful of small numbers, and then rebuilds the true result from that set of remainders. The number of multiplications grows only in proportion to the number of remainders. Scheme II was first posted as a preprint in April 2025 and appeared in a peer-reviewed journal in July 2026. It reached CUDA through a developer preview of version 13.4 in July and the release that followed in September.
NVIDIA's release notes say CUDA 13.4 picks Ozaki II automatically whenever it beats Ozaki I. On a B200, emulated double-precision matrix multiplication (DGEMM) reaches up to 175 TFLOPS, and the complex-number version (ZGEMM) up to 295 TFLOPS. Peak against peak, that is a little over four times what the B200's own double-precision circuits can do. For the next-generation Rubin, NVIDIA lists up to 212 TFLOPS for DGEMM, though the university notes that Rubin support in CUDA 13.4 is a developer preview.
No new hardware required
Scheme II needs no new circuitry. According to NVIDIA, it runs on Ampere-generation GPUs and newer. In a purely double-precision workload the AI units sit idle, and Scheme II puts them to work. Even the RTX PRO 6000 Blackwell Server Edition, which is not NVIDIA's top data-center part, is listed at up to 45 TFLOPS of emulated DGEMM.
National programs are paying attention. According to the university, RIKEN said in a March 2026 document that it plans to use this kind of emulation in the numerical libraries for FugakuNEXT, the successor to Japan's Fugaku supercomputer. The university also points to a February 2026 report in the trade outlet HPCwire that the US Department of Energy's Genesis Mission expects to lean heavily on the Ozaki scheme for its double-precision needs.
It is not a cure-all
First, the scope is narrow. As the university itself explains, the method currently targets matrix multiplication; NVIDIA has only begun bringing it to higher-level libraries for linear systems and eigenvalue problems.
AMD is openly skeptical. Speaking to HPCwire in March 2026, AMD Fellow Nick Malaya laid out the objections. The Ozaki scheme is not IEEE-compliant and does not give the same answer as FP64 hardware. Matrices whose elements differ by several orders of magnitude cause accuracy problems. With non-square matrices, it falls below native FP64 speed. And fewer than 10% of real-world HPC applications have been restructured around matrix multiplication in a way that would let them benefit. AMD is betting the other way with its science-oriented Instinct MI430X. HPCwire reported in August 2026 that the chip is expected to deliver 288 TFLOPS of native FP64 when it ships in early 2027, nearly nine times the 33 TFLOPS of NVIDIA's Rubin.
NVIDIA has built in some safeguards. According to the documentation for cuBLAS, CUDA's matrix library, the emulation estimates the bits it needs and falls back to native FP64 when that number exceeds a set limit. On a chip like the B300, though, that fallback path is itself the slow one. The same documentation notes that emulation can use up to 8 GB of working memory, and that infinities and NaNs may be handled differently from standard arithmetic.
Satoshi Matsuoka, director of the RIKEN Center for Computational Science, argued in a paper posted in May 2026 that the field should stop treating dedicated FP64 circuits as a requirement for scientific computing. But the paper has not been peer-reviewed, and Matsuoka himself notes that many of its performance figures are projections from a model, not measurements. Nobody can yet say that double-precision hardware has become unnecessary.
When a university's math enters the world's toolbox
CUDA, one of the most widely used platforms for GPU computing, now carries an algorithm born at a Japanese university. The method did not grow up inside one country, though. The original 2012 paper had three researchers in Japan and one in Germany: Siegfried Rump of Hamburg University of Technology. In November 2025, NVIDIA engineers published their own extension of the Ozaki scheme that automatically guarantees accuracy. Even Malaya, the critic, said AMD will support Ozaki emulation on its own chips. Once published, the math has been refined and challenged across countries and companies.
The libraries in your own lab or company probably carry equations from someone in a faraway country too. Have you ever looked up whose they are?
References
- https://www.shibaura-it.ac.jp/headline/detail/20260930_7070_51_1.html
- https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html
- https://docs.nvidia.com/cuda/cublas/index.html
- https://arxiv.org/abs/2504.08009
- https://link.springer.com/article/10.1007/s11075-011-9478-1
- https://arxiv.org/abs/2511.13778
- https://arxiv.org/abs/2606.06510
- https://www.hpcwire.com/2026/03/13/amd-hints-at-big-fp64-increases-in-mi430x-gpu-as-ozaki-underwelms/
- https://www.hpcwire.com/2026/08/03/amds-fp64-boost-with-mi430x-is-even-bigger-than-expected/
- https://eetimes.itmedia.co.jp/ee/articles/2610/02/news019.html
- https://developer.nvidia.com/cuda-toolkit-archive
Global Discussion
0 comments