Skip to content
View sunnydubey1111's full-sized avatar
💌
...cogito ergo sum...
💌
...cogito ergo sum...

Block or report sunnydubey1111

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sunnydubey1111/README.md

Sunny Dubey — Enterprise Java, Architecture, AI Engineering, AI Research

Enterprise Java, Spring Boot, microservices · Event-driven and distributed system design · RAG, agents, evaluations, and guardrails · Independent research on LLM agent reliability · Pune, India

☕ Senior Java Engineering Professional | 🏗️ Transitioning into Software Architecture | 🤖 AI Engineering | 🔬 AI Research

Senior Java engineering professional transitioning toward Software Architecture, with AI engineering skills and AI research experience.· 📍 Pune, India

⌨️ What I Do

  • Design and contribute to scalable, secure, and observable enterprise applications and distributed systems
  • Design microservices, APIs, event-driven architectures, and enterprise integration solutions
  • Modernize Java platforms using Spring Boot, resilient architecture, and cloud-native patterns
  • Apply security, performance, scalability, and reliability practices to production systems
  • Build AI applications using RAG, agentic workflows, multi-agent systems, evaluations, and guardrails
  • Work with context engineering, MCP, local LLMs, fine-tuning, and production AI deployment
  • Conduct independent research on the reliability, monitoring, and repair of LLM agents
  • Lead technical design, code reviews, stakeholder collaboration, troubleshooting, and root-cause analysis

🔬 AI Research

Real-Time Detection and Repair of LLM Agent Failures

arXiv DOI

Stars Forks Issues Last commit Top language Repo size Commit activity

A lightweight research system for detecting when an autonomous LLM agent begins to derail during execution and initiating repair before the failure propagates.

  • Learns normal execution dynamics from healthy agent trajectories
  • Combines reservoir computing, anomaly detection, and sequential monitoring
  • Uses behavioral, semantic, uncertainty, metadata, and grounding signals
  • Designed for real-time, CPU-only monitoring at approximately 200 µs per agent step
  • Evaluated across multiple agent frameworks, models, tasks, and failure types

🤖 AI Engineering

I am extending my Enterprise Java engineering background into production-oriented AI engineering through hands-on learning, projects, and independent research across:

  • LLM fundamentals, transformers, embeddings, and semantic similarity
  • Retrieval-Augmented Generation: vector, hybrid, SQL, graph, and multimodal RAG
  • Vector databases, document parsing, chunking, reranking, and semantic search
  • Tool-calling agents, ReAct workflows, memory, routing, and multi-agent orchestration
  • LangChain and LangGraph application development
  • RAG and agent evaluation, guardrails, observability, and AI security
  • Context engineering and Model Context Protocol (MCP)
  • Fine-tuning with LoRA/QLoRA, synthetic-data pipelines, SLMs, and local inference
  • FastAPI-based services, model serving, caching, routing, scaling, and cost optimization
  • Independent research in lightweight runtime monitoring and repair of LLM-agent failures

🧰 Tech Stack

Enterprise Java & Distributed Systems

Java Spring Spring Boot Spring MVC Spring Security Spring Data JPA Spring Cloud Hibernate JDBC Microservices REST SOAP OpenAPI OAuth 2.0 Spring Vault Feign

Distributed Systems & Integration

Apache Kafka JMS TIBCO EMS Eureka Hystrix ZooKeeper SAGA CQRS Event-Driven Architecture

AI Engineering & LLM Systems

Python FastAPI Streamlit LangChain LangGraph DSPy Qdrant Docling Semantic Router MCP LangSmith RAG AI Agents Multi-Agent Systems Evaluations Guardrails LoRA Unsloth Ollama vLLM Multimodal AI Reservoir Computing

Cloud, Platform & Delivery

AWS AWS AgentCore Azure AI Docker Kubernetes OpenShift GCP Jenkins Maven Bitbucket Jira Confluence

Cloud and container technologies above include professional exposure and collaboration with DevOps teams; they are not presented as my primary specialization.

Observability, Security & Quality

Kibana Zipkin Spring Actuator Wavefront JConsole VisualVM SonarQube Checkmarx Black Duck JUnit Mockito PowerMockito Log4j2

Data & Runtime Platforms

Oracle DB2 SQL Server MongoDB MySQL WebSphere Tomcat JBoss WebLogic

Supporting Technologies

Angular AngularJS HTML5 XML JSON Swagger UML Eclipse IDE Spring Tool Suite IBM RAD


🧭 Architecture Focus

  • Distributed systems and microservices
  • API and enterprise integration architecture
  • Event-driven architecture and asynchronous messaging
  • Resilience, fault tolerance, and graceful degradation
  • Application security, authentication, and secrets management
  • Performance engineering, scalability, and database optimization
  • Observability, runtime monitoring, and distributed tracing
  • Legacy modernization and cloud adoption
  • Production-oriented design of AI-powered systems

🎓 Background

  • M.S. in Electrical Engineering — Texas A&M University–Kingsville, US
  • B.E. in Electronics and Telecommunication Engineering — SRTMU, India
  • Certificate Programme in Project Management — IIM Indore, India

📝 Engineering Philosophy

  • Architecture should simplify change, not merely organize complexity
  • Production reliability matters more than impressive prototypes
  • AI systems need evaluation, observability, security, and failure handling
  • Every component should justify its operational cost
  • Good systems are secure, explainable, maintainable, and resilient by design

📊 GitHub Activity

Contribution streak: total contributions, current streak, and longest streak

Contribution activity over the last 31 days

Repositories per language Most committed language

Overall GitHub stats Most productive time of day, IST


🤝 Connect

LinkedIn Email GitHub


“Years of engineering teach you how to build software.
Architecture teaches you how to make the right systems.
AI is changing what those systems can become.”

Pinned Loading

  1. agent-trajectory-sentinel agent-trajectory-sentinel Public

    Real-time detection and repair of LLM agent failures — a one-class behavioural monitor at ~200 µs/step, with 2,823 committed traces.

    Python 5