Dheemai
VivaGuard · Deployment Options

VivaGuard — Choosing How to Run It

Three ways to deploy VivaGuard voice-integrity for online exams & interviews.

Deployment option Tech spec1 Where your data lives Who runs it Time to launch Scaling Available today
On-premise
in your own data centre
1 server (4–8 vCPU, 8–16 GB) for a pilot; add a GPU box (1–2× NVIDIA L4) for exam-mode at full scale. Docker-based; can run fully offline / air-gapped. Entirely inside your network — nothing ever leaves your servers. Simplest path for DPDP / biometric compliance. You provide and run the hardware, updates and backups. Longest — hardware, install, and your security review. Limited by the hardware you provision. ✅ Ready to install now
In your own Google Cloud Runs in your GCP project (Cloud Run or GKE): a 4 vCPU / 8 GB API tier plus a 2–4× NVIDIA L4 GPU pool for exam-mode; voiceprints in your Cloud SQL or Cloud Storage. Stays in your own cloud project and region — you keep full control and data residency. Your cloud team runs it; we help you deploy and can co-manage. Moderate — a cloud deployment into your project. Elastic — scales up and down automatically. ✅ Ready (with a small storage add-on for scale)
Fully managed by Dheemai Runs on Dheemai's autoscaling GKE + GPU pool; you simply call the API with your key. Candidate audio is processed on our infrastructure — needs a data-processing agreement and explicit consent. We run everything — scaling, updates, monitoring. You manage nothing. Fastest — live as soon as we issue your API key. Elastic — we scale it for you. 🔜 On our roadmap (multi-tenant service in development)
Footnotes & assumptions
  1. Tech spec is sized for up to ~1,000 concurrent exam sessions at a peak of roughly 10–30 exam-mode analyses per second. A small pilot (a few dozen users) runs comfortably on a single CPU server with no GPU.
  2. Exam-mode (analysing a spoken answer for reading vs. genuine speech) is the compute-heavy part — ~60 seconds of CPU per 60-second clip, or ~1–2 seconds on a GPU. Identity verification is light (~1–2 seconds) and needs no GPU.
  3. GPU figures are based on NVIDIA L4-class accelerators; counts scale roughly linearly with peak exam-analyses-per-second.
  4. “Available today” reflects the current single-tenant product. The fully-managed, multi-tenant service requires additional engineering before launch.
  5. DPDP = India's Digital Personal Data Protection Act; a person's voice is treated as sensitive / biometric personal data, which is why data location matters.
Our recommendation: for production with real candidate voices, choose On-premise or In your own Google Cloud — your data stays under your control and both are available today. Use the fully-managed option for a quick, consented pilot while you evaluate.