Bhanu Pallakonda

Bhanu Pallakonda

I post-train large language models and engineer the agent harnesses around them — AI that runs from cloud GPU clusters down to disconnected, compute-constrained edge sites.

Lead AI Engineer, Armada — Seattle, WA

3 granted US patents · 3 peer-reviewed papers · 60+ citations

About

I'm the first AI hire at Armada, where I lead model post-training (SFT and RL) and agent-harness engineering for a platform that operates conversational and agentic AI at local, bandwidth-limited, and disconnected sites.

My work covers the full lifecycle: fine-tuning open-weight models, quantizing and serving them from H100 clusters down to CPU-only edge nodes, and building the multi-agent and skill-agent systems that make them dependable. I'm a named inventor on three granted U.S. patents — all commercialized in Armada's shipping product — and I publish on edge language models, AI security, and geospatial ML.

Away from the keyboard: cricket, films, and long road drives.

Experience

Lead AI Engineer, Armada

2023 — present

Seattle, WA · first AI hire

I own the AI stack for Armada's edge platform — model post-training and agent-harness engineering, from NVIDIA H100/A100 GPUs down to CPU-only nodes. The work has moved through three generations:

Conversational AI assistant
Built the platform's core assistant from scratch on open-weight Llama models with an LLM-as-router design; deployed end-to-end on AKS with vLLM, and quantized for on-device edge inference.
Multi-agent systems
Designed a production supervisor–sub-agent system with modular inference routing; hardened against prompt injection, with observability and evaluation built in.
Skill-agent harness & coding agents
Re-architected the assistant as a composable skill + coding agent — deep agents, code sandboxing, and Armada-specific asset skills — alongside AI-safety research on backdoors in tool-using LLMs.

Machine Learning Engineer, Fincare Small Finance Bank

2019 — 2020

BERT-based banking chatbot (intent and entity extraction), ID-card detection, and privacy-preserving field-masking pipelines in production.

Research Assistant (NSF), Texas A&M · ML Intern, Productiv

2022

Pancreatic-cancer prediction on clinical datasets; contract field-extraction with fine-tuned LayoutLMv3.

Selected systems

Production AI systems designed and shipped at Armada. Source is proprietary.

Edge conversational AI assistant

Armada's core assistant, built from scratch on open-weight Llama models behind an LLM-as-router intent layer. Serving is tiered: vLLM on H100/A100 nodes, quantized GGUF weights via llama.cpp on CPU-only and disconnected sites — deployed end-to-end on AKS.

Multi-agent system

A supervisor–sub-agent architecture for task decomposition, routing between reasoning and non-reasoning models across API and self-hosted endpoints. Hardened against prompt injection; observed with Langfuse tracing, an online accuracy evaluator, Grafana alerting, and CI-integrated agentic regression tests.

Skill-agent harness & coding agents

The current agentic layer: deep agents with code sandboxing and Armada-specific asset skills. The assistant plans, writes, and executes sandboxed code against domain assets as a composable skill + coding agent.

Embeddings & retrieval service

A production embeddings microservice powering retrieval-augmented generation across the platform — vector embeddings and semantic search that ground the assistant and agents in domain data.

Publications

Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs

arXiv:2603.03371

Temporal backdoors in tool-using LLMs: latent triggers that surface malicious behavior only under specific conditions, studied through RL fine-tuning and distributed multi-GPU experiments. To be submitted to NeurIPS 2026.

Camera Control at the Edge with Language Models for Scene Understanding

doi:10.1109/ICCAR…3044

OPUS, an LLM framework for PTZ camera control with contextual scene understanding — +35% over advanced prompting and +20% task accuracy against closed-source models, via supervised fine-tuning on synthetic data.

Mapping the Storm: Geospatial Impacts of Severe Weather on LEO Network Performance

doi:10.1145/3748…2720

A geospatial analysis of how severe weather degrades LEO satellite-network performance, informing connectivity resilience for remote infrastructure.

Patents

All titled “Edge Computing Units for Operating Conversational Tools at Local Sites” — one continuation family, assigned to Armada Systems and commercialized in the shipping product.

US 11,960,515

granted Apr 2024 · 34 citations · patent ↗

Lead patent of the family.

US 12,001,463

granted Jun 2024 · 25 citations · patent ↗

US App. 18/810,025

pending

Skills

post-training
SFT, RLHF / RL, model evaluation, red-teaming, quantization (GGUF), synthetic data
agentic systems
multi-agent, skill & coding agents, LangChain / LangGraph, deep agents, code sandboxing, RAG, prompt-injection defense
serving & infra
vLLM, llama.cpp, Azure / AKS, Docker, Kubernetes, CI/CD, Langfuse, Grafana, Redis, Celery
languages
Python, PyTorch, TensorFlow, FastAPI, JavaScript, React, SQL

Education

2021 — 2023
M.S. in Computer Science — Texas A&M University
2015 — 2019
B.Tech. in Electrical Engineering — Indian Institute of Technology (IIT) Tirupati

Contact

Open to research collaborations and hard problems in applied LLMs. The fastest way to reach me is email.