Skip to content
View punith1006's full-sized avatar
  • Global Knowledge technologies
  • Bangalore
  • 12:13 (UTC -12:00)
  • X @vs_punith

Block or report punith1006

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
punith1006/README.md

はじめまして、プニートです 👋

Applied AI Engineer & Solutions Architect

Hey there! I'm an engineer based in Bangalore working across Applied AI, systems architecture, and backend engineering. I am drawn to hard, high-leverage technical problems—the kind where standard tutorials run out and you have to build the path forward from first principles. I am interested in collaborating with high-velocity teams, technical agencies, and founders tackling frontier challenges who need engineers capable of taking an ambiguous problem, mastering the necessary tech quickly, and delivering dependable software.

Proof in Production (What I've Built)

  • Sovereign GPU Cloud Infrastructure (LaaS):
    Served as Solutions Architect and Lead Engineer for a bare-metal, multi-tenant private GPU cloud platform. Implemented hardware-level fractional virtualization (CUDA MPS, HAMi-core), low-latency (<25ms) WebRTC desktop streaming via NVENC, and persistent ZFS storage—multiplying single-GPU seat density by 4x–8x and slashing infrastructure costs by ~80% compared to commercial public cloud instances.

  • Autonomous Multi-Agent Workflows (AIA):
    Architected an automated B2B customer acquisition pipeline for SMBs using the Google Agent Development Kit (ADK). Orchestrated multi-agent chains that conduct autonomous deep-web research, brand positioning analysis, competitor benchmarking, and demographic extraction to synthesize high-precision Ideal Customer Profiles (ICPs) and tailored outreach.

  • High-Throughput LLM Serving & Inference Engine: Setup an optimized local model serving infrastructure using vLLM and SGLang, going well beyond stock deployments. Understood the engine internals in detail and configured PagedAttention KV-cache management, continuous batching for high concurrency, and intelligent request routing for optimized inference—achieving a ~40% throughput increase and significant token latency reductions over baseline serving setups.

🤝 Let's Connect

If your team or agency is building something ambitious and needs an engineer who moves fast, adapts without friction, and builds with rigor, feel free to reach out.

Pinned Loading

  1. LaaS-Strata LaaS-Strata Public

    Sovereign private GPU cloud platform. Slices physical GPUs into 4–8 isolated, 60 FPS in-browser AI workstations streaming full graphical Linux desktops directly to any web browser with persistent s…

    TypeScript

  2. aia aia Public

    Python

  3. gulfood_replit gulfood_replit Public

    TypeScript

  4. moviecompanion moviecompanion Public

    Python

  5. analytics_agent analytics_agent Public

    Python