On the same GPU, what wall-time cost is required to reach a given numerical error with Ramanujan CUDA versus simulated Quantum Amplitude Estimation?
Compare Ramanujan's 1914 pi series in a CUDA C++ kernel with Quantum Amplitude Estimation in CUDA-Q/cuQuantum, validated against exact SymPy ground truth.
On an RTX 5070 Laptop GPU, the CUDA implementation reached double-precision saturation (~16 digits) in ~2.6 ms, while simulated QAE reached ~5 digits in ~0.44 s. No crossover was observed on this hardware.
Run the optional H100 axis, then open the hardware-agnostic benchmark format to community submissions.