Workhorse Inference. Predictable Scale.
Maximum FLOPs per rupee. The default compute node for most AI workloads.

Most enterprise AI workloads do not need flagship GPUs. They need reliable, high-throughput inference at a price that scales linearly. JOHNAIC 32 is the workhorse node: two RTX 5070 Ti GPUs, 32 GB of total GPU memory, and the same validated software image as every other node in the platform.
Low VRAM per GPU is accepted in exchange for high FLOPs per rupee. Panini's model routing compensates for single-GPU memory limits.
Add capacity in ₹10L increments
No re-architecture. No procurement theater. Drop in another J32 and the cluster grows.
Same image, same support playbook
Pre-loaded with Multix + Titan + Panini. Validated, burned in, and ready to serve models.
100 GbE fabric as standard
Dual 100 GbE is not an upgrade. Every J32 participates fully in the cluster fabric from day one.
- —Document QA and RAG pipelines
- —Vision-language tasks: OCR, form extraction, image understanding
- —Embedding and retriever models
- —High-concurrency inference with vLLM / PagedAttention
- —Control plane for small clusters (when J0 is not present)
- —70B+ parameter models — use JOHNAIC 192
- —Standalone evaluation where a single J32 is sufficient
- —Small models: 7B–14B parameters (ideal, runs on a single GPU)
- —Medium models: 32B parameters (tight; tensor parallelism or quantization recommended)
- —Large models: 70B+ parameters — does not fit. Use J192.
- —Document QA / RAG
- —Vision-language (OCR, form extraction)
- —Embeddings / retriever models
- —High-concurrency inference
Validate J32 on your workload.
Run your documents, models, and inference patterns on a J32 sandbox inside your perimeter. No cloud, no metered API.