Arjun GaneshGoverned AI · Distributed systems
Colour theme
CV
04

Research

Measured investigations with explicit questions, methods, current results, and unresolved work.

01

q1729

active
Question

On the same GPU, what wall-time cost is required to reach a given numerical error with Ramanujan CUDA versus simulated Quantum Amplitude Estimation?

Method

Compare Ramanujan's 1914 pi series in a CUDA C++ kernel with Quantum Amplitude Estimation in CUDA-Q/cuQuantum, validated against exact SymPy ground truth.

Current result

On an RTX 5070 Laptop GPU, the CUDA implementation reached double-precision saturation (~16 digits) in ~2.6 ms, while simulated QAE reached ~5 digits in ~0.44 s. No crossover was observed on this hardware.

Next

Run the optional H100 axis, then open the hardware-agnostic benchmark format to community submissions.

CUDA C++CUDA-QcuQuantumSymPyNIM / Nemotron
02

llm-qlab

active
Question

How do GGUF quantization and CPU/GPU layer placement change inference throughput and memory use on consumer NVIDIA hardware?

Method

Separate prefill and decode with llama.cpp counters; discard warmups; report repeated-run variance; verify memory clock state; reject unstable or paged runs.

Current result

Across three 7B model families, decode throughput fell monotonically as quantized weight size increased. Full Llama-2 offload measured 6.2x CPU-only decode throughput.

Next

Add batch-size and context-length sweeps, quality regression by quantization format, and a datacenter-GPU comparison.

Pythonllama.cppCUDAGGUF
Question

Where do GPU implementations outperform CPU baselines across common algorithms and input sizes?

Method

Run configurable CPU/GPU benchmark sweeps, export the measurements to CSV, and inspect empirical complexity and speedup in an interactive dashboard.

Current result

In the published sweep, radix sort is the strongest GPU result, while BFS and reductions remain slower than their CPU baselines at the tested sizes.

Next

Extend the documented benchmark protocol to additional algorithms, input shapes, and verified GPU runs.

PythonCuPyNumba CUDADash
Learning

Technical notebooks

Microsoft IQiq-series ↗Completed Microsoft IQ learning cookbooks with executed Foundry IQ notebook outputs.