Nvidia H1009 min read
How to Install FlashAttention-3 on NVIDIA Hopper GPUs
Learn how to install Flash Attention 3 on NVIDIA Hopper GPUs, match CUDA and PyTorch, build from source, verify the package, and fix common errors.
Kolkata, India
I design agentic workflows, production RAG pipelines, and LLMOps infrastructure enterprise AI systems built to hold up under real traffic, real users, and real compliance pressure.
Open to
Planner agent
LangGraphBreaks the request into steps and routes work to specialists
RAG retrieval
Azure AI Search · Pinecone
MCP tool calls
JIRA · SharePoint · Web
Evaluation gate
Human-in-the-loopRAGAS scoring and an approval checkpoint before anything moves forward
Ship to production
Monitored, logged, and easy for teams to run
The agentic pattern behind my recent builds — planning, grounded retrieval, tool execution, and a human gate before production.
From the blog
Nvidia H1009 min read
Learn how to install Flash Attention 3 on NVIDIA Hopper GPUs, match CUDA and PyTorch, build from source, verify the package, and fix common errors.
Hugging Face Peft8 min read
Learn how to use QLoRA Hugging Face to fine-tune Llama 3 in 4-bit NF4 with PEFT and TRL, then compare VRAM use against LoRA and DoRA adapters.
Flashattention-28 min read
Learn how Flash Attention vs grouped query attention affects latency, VRAM, and KV cache size, with PyTorch benchmarks and implementation guidance.
Vllm8 min read
Learn how to implement paged attention in PyTorch with a block KV cache, decoding tests, memory math, benchmarks, and a clear mapping to vLLM.
Trusted by teams
High-compliance AI systems for regulated sectors — global banking, healthcare, enterprise platforms, and AI-native products — where "it mostly works" isn't good enough.
Where it shows up
The common thread is building systems people can actually use day to day, not just demo once and forget.
Brand platforms delivered through employer and consulting engagements at TEKsystems, Syren Technologies, RWS Moravia, and Capgemini, alongside direct product work.
Delivery record
Every figure below is sourced from a specific role or system in the experience timeline further down this page.
Systems delivered
9+
Enterprise AI systems shipped end to end
Daily queries
10K+
Scaled in production on a supply-chain copilot
Production uptime
99.9%
Maintained across live enterprise workloads
Efficiency gain
67%
Reduction in manual operational workloads
Certifications
6
Verified cloud, data, and AI credentials
Enterprise reach
10+
Organizations and teams supported
Professional experience
Across 5+ years
The work moved from data platforms into search, copilots, and enterprise GenAI systems, with reliability and delivery pressure increasing at each step.
TEKsystems Global Services Pvt. Ltd
Cut intake-to-ticket turnaround from 2+ days to under 10 minutes, compressing P95 pipeline latency by 85% with parallelized Redis-backed workers
Improved learner skill-gap detection accuracy by 37% across 6K+ monthly sessions with adaptive quiz generation and continuous RAGAS evaluation
Syren Technologies Private Limited
Scaled a supply chain AI copilot to 10K+ daily queries, reducing latency by 63% while maintaining 99.9% uptime
Boosted contextual recall by 34%, improved multi-turn response accuracy by 28%, and reduced context errors by 42%
RWS Moravia India Private Limited
Launched the company's first production-grade multilingual enterprise RAG platform and AI agent
Cut manual knowledge lookup by 43% and established rigorous BLEU and NDCG evaluation benchmarks
BinaryERP Software Private Limited
Built semantic search prototypes that drove adoption across 13 B2B users
Improved dataset usability by 26% and laid the foundation for enterprise NLP initiatives
Capgemini Technology Services India Limited
Reduced nightly batch runtime by 95% for a global banking client
Decreased batch failures by 61% while preserving SLA compliance and improving ETL validation automation
Enterprise AI deployments
What they have in common
Every system here started from a clear user task, shipped with metrics tied to adoption or efficiency, and held up under daily production use.
Multi-agent automation with JIRA, SharePoint & live web search
Adaptive sales enablement for regulated medical teams
Voice-first analytics for teams that live in data
Interview recordings turned into clean question lists
A knowledge copilot for documents, tables, and daily decisions
AI content intelligence for building a credible LinkedIn voice
Recruiter intelligence for faster resume screening
A private document assistant for faster legal review workflows
One assistant for HR, finance, policy, and payroll questions
Where I help
What usually matters
Some teams need a stronger retrieval core, others need multi-agent workflows or a safer path from pilot to platform. These are the patterns I most often design around.
Design multi-agent systems that coordinate tools, memory, approvals, and specialist roles without becoming hard to manage.
Best suited for
Teams replacing repetitive manual work with planner, reviewer, and executor workflows.
Clear agent roles and handoffs
Human review steps and safety rails
Build retrieval systems that surface the right context quickly, keep answers grounded, and make knowledge easier to use.
Best suited for
Products and internal copilots that need trusted answers from documents, data systems, and live business context.
Ingestion, chunking, and indexing plan
Hybrid retrieval with reranking and citations
Improve quality, speed, and consistency with the right mix of fine-tuning, prompt shaping, dataset design, and inference optimization.
Best suited for
Teams that already see promise but need better domain accuracy, tone control, or task fit.
Dataset curation and benchmark framing
LoRA or PEFT fine-tuning workflow design
Shape the end-to-end AI setup so it is secure, observable, and easy for real teams to run over time.
Best suited for
Teams moving from early AI experiments to dependable platforms, governance, and long-term use.
A reference setup for services and data flow
Security, compliance, and deployment patterns
Engagement models
Every engagement starts from one of these formats. Baseline rates and a delivery estimate are shared in the first scoping call.
1-2 weeks · Fixed scope
A structured review of an existing AI system: retrieval quality, agent design, evaluation coverage, cost, and reliability risks.
You get
Scored findings report with a prioritized remediation roadmap.
2-3 weeks · Fixed scope
Shape an AI idea into a validated plan: use-case framing, data readiness, architecture options, and a working proof of concept.
You get
Reference architecture, PoC, and a delivery estimate.
4-12 weeks · Milestone-based
Design and ship a production AI system: agent workflows, RAG pipelines, evaluation loops, deployment, and team handoff.
You get
Production system with monitoring, documentation, and handoff.
Monthly retainer
Ongoing architecture ownership for teams that need senior AI leadership without a full-time hire: reviews, roadmaps, delivery oversight.
You get
Weekly architecture sessions plus async design reviews.
Production tech stack
What ties it together
The tools change depending on the problem, but the goal stays the same: grounded answers, reliable workflows, and systems teams can keep using after launch.
Building end-to-end LLM systems, agent workflows, and prompt strategies teams can rely on.
Often includes
Building AI solutions on Azure with cloud services, search, deployment patterns, and data tooling.
Often includes
Measuring retrieval and answer quality so AI systems improve with clear feedback loops.
Often includes
Building retrieval layers with semantic indexing, hybrid search, and embedding pipelines for fast, accurate lookup.
Often includes
Keeping AI applications dependable with deployment workflows, monitoring, orchestration, and automation.
Often includes
Strong foundations in Python, ML tooling, and data platforms that support dependable delivery.
Often includes
Credentials
Built around practice
A focused mix of cloud, data, and AI credentials that support the delivery side of the work.

Core Azure services, security, pricing, and cloud basics.

Azure data services, storage models, and analytics foundations.

Lakehouse concepts, workspace workflows, and platform essentials.

LLM concepts, prompt basics, and responsible GenAI grounding.

Secure, scalable Azure Databricks architecture and platform design.
DAG design, scheduling, and workflow orchestration fundamentals.
Where it started
Maulana Abul Kalam Azad University of Technology
How I work
Why I work this way
I've watched too many AI projects die in a six-month build phase nobody saw. So every step here produces something you can see and test — and the system gets stronger every week.
I start by understanding your workflow, your users, and your data — and what success would actually look like. No code until that's clear.
Then I pick the retrieval strategy, agent topology, models, and guardrails that fit your constraints — not whatever is trending that week.
I ship a working end-to-end slice early, then iterate against real data and real user feedback. You see progress every week, not at the end.
Before launch I build eval pipelines, tune latency and cost, and make sure failures show up on a dashboard — not in a user complaint.
I deploy with monitoring in place, write the runbooks, and make sure your team can run the system confidently without me.
Professional profile
I'm an enterprise AI architect based in Kolkata, India. For 5+ years I've worked across data engineering, search, and generative AI, building systems where answer quality, reliability, and business value all matter at the same time.
What I focus on
As a Lead AI Architect, I care about AI systems that keep working once teams start relying on them. That mindset comes from my years in cloud data platforms, automation, and reliability-focused engineering.
Much of my recent work centers on LangGraph orchestration, RAG, and domain-specific AI products. Recent platforms have supported 10K+ daily queries while maintaining 99.9% uptime in live environments.
I engineer for outcomes teams can measure: lower latency, automated workflows, and faster decisions. That focus has delivered efficiency gains as high as 67% while keeping each system clear enough for teams to maintain after launch.
How I work
Strong AI systems earn trust through clear results and careful testing. I set up evaluation early so teams can see what is working and improve with confidence.
Retrieval quality is checked before scale becomes the priority.
Answer quality is reviewed with clear task-based checks.
Decisions stay tied to user needs and business goals.
The best AI products keep working when data is messy, traffic grows, and more than one team needs to use or support them.
Workflows are kept clear before complexity piles up.
Latency, fallbacks, and caching are planned from the start.
Monitoring and team handoff are considered alongside model behavior.
The most valuable AI work saves time, improves decisions, or removes repetitive effort. That is the standard I like to build toward.
The business goal stays visible from discovery through launch.
Security, guardrails, and handoff readiness are built in early.
The final system should feel easy to use, extend, and trust.
Quick answers
Something else on your mind?
If your question isn't here, just ask — a quick call or a short email works. I read everything myself and usually reply within a day.
Swarnava Dutta is a Lead AI Architect based in Kolkata, India, with 6 years of experience in generative AI engineering. He designs and ships production-grade AI systems — multi-agent workflows built with LangGraph, retrieval-augmented generation (RAG) architectures, and enterprise LLM platforms on Azure OpenAI and Azure AI Foundry. His most recent project is LindsAI, a multi-agent system with tool calling into JIRA and SharePoint plus dynamic web search. He currently works at TEKsystems as a Senior GenAI Engineer and Technical Lead.
Next step
Start with the resume, look through the work, or send a quick note first. Whatever helps you get context fastest.
Let's talk
Best starting points
Bring a consulting brief, a full-time opportunity, an architecture question, or a specific AI system that needs clearer shape.
Calendar not loading? Open the booking page in a new tab.
Project brief
Share the problem, timeline, or role you have in mind. A short brief is enough.