Skip to content
SHIVSASTRA
YS-V22026AI + Deterministic Systems
Core Principle: AI understands. Java decides.

A WhatsApp AI Agent for Welfare-Scheme Discovery.

Talk naturally. Let the agent extract your profile attributes from conversational messages. Deterministic Java rules evaluate eligibility over 82 normalized schemes.

ROLE

Architecture · Backend · Security · AI Integration

STACK

Java · Spring Boot · PostgreSQL · Groq

STATUS

Release Ready (Documented Conditions)

CHANNEL

WhatsApp + Portfolio Demo

01 // VERIFIED SYSTEM BENCHMARKSVERIFIED PROJECT METRICS
NORMALIZED SCHEMES
82

63 Central · 11 State · 8 Philanthropic

Seeded into 1NF relational schema on Supabase PostgreSQL 17.

AUTOMATED TESTS PASSING
42/42

42 Passing Tests Across Core Suites

Backend release-gate verification recorded 42/42 passing tests.

CRITICAL & HIGH FINDINGS
0 / 0

0 Critical / High Findings in Final Security Audit

Verified in release-gate security audit with zero high-severity findings.

INDEXED RULES EVALUATION
B-Tree

PostgreSQL Relational Indexing

Eligibility rules are evaluated through indexed relational criteria queries rather than unstructured prompt context or table scans.

Verification note: Benchmarks and tests reflect verified relational schema normalization, automated test suite passes, and security boundary assertions. External network delivery depends on telecommunication and messaging gateway factors.

02 — PRODUCT EVIDENCE

Real-World Multilingual Scheme Discovery

Historical interaction captures from the original Yojna Setu conversational pipeline. These unedited captures demonstrate how real citizens interface with natural-language intake, slot filling, and eligibility resolution.

EVIDENCE 01 // WHATSAPP CONVERSATIONAL INTERACTIONHISTORICAL PRODUCT ARTIFACT
Historical Yojna Setu WhatsApp interaction showing multilingual profile intake, missing-gender clarification and scheme recommendations.

Intake & Multi-Variable Demographic Extraction

The user triggers a session reset and inputs multi-variable demographic parameters in conversational Hinglish ('namaste, main ek student hoon, 20 saal, UP se, general hoon, income 1.5 lakh , hindu'). The engine extracts the attributes, identifies the missing gender slot, asks for clarification ('Aap purush hain ya mahila?'), and recommends 5 matched schemes upon receiving 'Purush'.

Historical Product Interaction — Captured during original prototype conversational testing.

EVIDENCE 02 // WHATSAPP CONVERSATIONAL INTERACTIONHISTORICAL PRODUCT ARTIFACT
Historical Yojna Setu WhatsApp interaction showing deadline retrieval and document requirements.

Temporal Deadlines & Prerequisite Documentation

The user inquires about deadlines by messaging 'Deadline', receiving upcoming dates for UP Mukhyamantri Abhyudaya Yojana. Following with 'Documents', the system delivers the exact documentation checklist required across matching schemes (Aadhaar Card, Ration Card, Income Certificate, Class 12 Marksheet, Admission Letter).

Historical Product Interaction — Captured during original prototype conversational testing.

CONVERSATIONAL WALKTHROUGH

The 3-Step WhatsApp Experience

The screenshots record an actual conversation. The user talks naturally in Hinglish, the bot asks for any missing profile detail, and the backend returns matching schemes with deadlines and documents:

  1. PHASE I: INTAKE & CLARIFICATION

    1. User sends Reset to start fresh.
    2. User provides unformatted Hinglish with age, state, student occupation, income, and religion.
    3. System identifies missing gender and asks for clarification (“Aap purush hain ya mahila?”).

  2. PHASE II: SCHEME EVALUATION

    4. User replies Purush.
    5. All demographic criteria are satisfied; rules engine filters 82 schemes down to 5 qualified programs.
    6. User receives verified scheme titles with direct portal application URLs.

  3. PHASE III: GUIDANCE & CHECKLISTS

    7. User sends Deadline, receiving upcoming cutoff dates.
    8. User sends Documents.
    9. System details the required identity and verification paperwork for each scheme.

Context: These captures show the front-facing WhatsApp conversation from prototype testing. In V2, the entire backend was rebuilt from scratch with the deterministic Java rules engine and security controls explained below.

03 — FORENSIC AUDIT & FIXES

From Prototype to Hardened Backend

The original prototype had critical security and correctness issues. Here are the 5 key engineering fixes made in V2 to harden the backend architecture.

01 //

Credential Security

BEFORE // LEGACY PROTOTYPE

Plaintext database credentials committed directly in startup shell scripts (start-app.sh).

AFTER // V2 HARDENED SYSTEM

Environment-based secrets with strict startup validation and complete Git history cleanup.

02 //

Eligibility Matching Logic

BEFORE // LEGACY PROTOTYPE

Brittle substring evaluation: scheme.getGender().contains("MALE") erroneously returned true for "FEMALE".

AFTER // V2 HARDENED SYSTEM

Deterministic relational rules (EligibilityEngine) using typed enums and database join criteria.

03 //

Media & SSRF Defense

BEFORE // LEGACY PROTOTYPE

Unvalidated remote media downloads allowing potential SSRF against private networks and metadata endpoints.

AFTER // V2 HARDENED SYSTEM

SSRF protection: HTTPS enforcement, Twilio host allowlist, private IP blocking, and bounded streaming.

04 //

Conversation State

BEFORE // LEGACY PROTOTYPE

In-memory ConcurrentHashMap causing state loss across server restarts and cross-thread concurrency issues.

AFTER // V2 HARDENED SYSTEM

Persistent database-backed conversation state machine (ConversationSession) with session recovery.

05 //

Webhook Idempotency

BEFORE // LEGACY PROTOTYPE

Retried provider webhooks generated duplicate outbound messages and corrupted multi-step state.

AFTER // V2 HARDENED SYSTEM

Idempotency layer: in-memory cache backed by unique constraints on webhook_events in PostgreSQL.

Security Scope: Hardening focused on verifiable engineering controls: eliminating hardcoded secrets, guaranteeing mathematical determinism in benefits matching, sanitizing inputs and remote media, and enforcing strict session idempotency.

04 — SYSTEM ARCHITECTURE

How the System Works

The main design principle is simple: AI extracts information. Java rules decide eligibility. The language model handles informal, multilingual chat, while the deterministic Java backend enforces statutory rules.

STEP 01

User Ingress

WhatsApp & Browser Demo

Users send conversational messages in Hindi, Hinglish, or English, or spoken voice notes.

Inbound requests enter through the Twilio webhook adapter or the local portfolio demo adapter.
STEP 02

Webhook Security

Spring Security Layer

Checks webhook signatures, verifies idempotency, and filters external media downloads.

Rejects forged payloads and validates media URLs before anything is processed.
STEP 03

AI Profile Extraction

Groq Cloud (LLM + Whisper)

Parses conversational text and voice into structured profile fields (age, state, income, caste).

Bounded role: The AI only extracts demographic parameters. It has zero authority over scheme eligibility.
STEP 04DECISION MAKER

Java Eligibility Rules

Deterministic Engine

Evaluates exact statutory rules: age limits, income ceilings, caste, and state residency criteria.

The sole authority for qualification decisions. Replaces probabilistic text matching with deterministic relational rules.
STEP 05

PostgreSQL Database

Supabase (yojna_setu Schema)

Stores 82 normalized schemes, criteria tables, conversation state, and webhook deduplication records.

Indexed relational queries evaluate criteria directly against structured B-Tree indices.
Pipeline Flow:Natural-language intake→Structured criteria→Deterministic Java rules→PostgreSQL
B-Tree Indexed Rules Evaluation
05 — CORE DESIGN PRINCIPLE

AI Understands. Rules Decide.

Even if AI extraction is imperfect, it cannot directly decide eligibility. The final decision is made by deterministic Java rules evaluated against verified government scheme criteria.

Where AI Is Used

Language models are great at parsing messy, conversational inputs. In Yojna Setu, Groq Cloud (GPT-OSS-20B and Whisper Large V3 Turbo) is used strictly for:

  • 01.Understanding Natural Language: Handling informal messages in Hindi, Hinglish, and English without requiring rigid menus.
  • 02.Profile Extraction: Pulling key details like age, state, annual income, caste, and occupation into structured fields.
  • 03.Voice-to-Text: Transcribing WhatsApp audio voice notes into clean text for downstream extraction.

Where AI Is NOT Used

The AI has zero authority over scheme eligibility decisions. All qualification is handled by the Java rules engine because:

  • 01.Bounded AI Impact: The AI only extracts structured profile information. It does not decide whether a user qualifies for a scheme.
  • 02.Clear Audit Trail: Every matched scheme returns the exact mathematical reason (e.g., age ≤ 25, income ≤ ₹2.5L, state = UP).
  • 03.Reliable Testing: Rules can be tested with unit tests and database queries rather than hoping prompts don't drift.

Interview Summary: If the user enters a typo or ambiguous input, the AI might misinterpret a demographic slot, but the Java engine will only evaluate what is structured. The AI cannot fabricate an entitlement or bypass statutory income limits.

06 — SECURITY ENGINEERING

Security Controls

Key security controls implemented in V2. Instead of generic marketing claims, these represent concrete software defenses built into the backend.

CONTROL 01

Webhook Verification

Twilio webhook signatures are checked before processing requests.

Inbound requests without valid HMAC signatures are immediately dropped, preventing forged or spoofed messages.

CONTROL 02

SSRF Protection

External media URLs are validated before the server downloads them.

Restricts media downloads to verified Twilio hosts, blocks private IP ranges and cloud metadata endpoints, and enforces 5MB stream limits.

CONTROL 03

PII-Safe Logging

Sensitive phone numbers and demographic data are masked in logs.

Phone numbers are stored as salted SHA-256 hashes, and a custom Logback converter masks mobile numbers in application logs.

CONTROL 04

Webhook Idempotency

Repeated webhook deliveries do not create duplicate processing.

Unique database constraints on webhook event IDs prevent re-processing retried requests or sending duplicate replies.

CONTROL 05

Persistent Conversation State

Conversation state is stored so multi-step conversations survive restarts.

Multi-turn demographic collection is backed by PostgreSQL sessions rather than volatile in-memory maps.

RELEASE GATE STATUS

0 Critical / High Findings

Verified in the final independent security audit before release.

Backend test suites verify webhook HMAC, SSRF revalidation, and session isolation.

07 — AUTOMATED TEST SUITE

42/42 Tests Passing: Core Test Areas

Backend release-gate verification recorded 42/42 passing tests across four primary areas. Tests were written around actual regression vectors from the prototype.

AREA 016 TESTS

Eligibility Rules

Deterministic age bounds, income ceilings, gender matching, and state residency rules.

STATUS✓ ALL PASSING
AREA 027 TESTS

Webhook Security

HMAC signature verification, replay protection, and provider response formatting.

STATUS✓ ALL PASSING
AREA 0316 TESTS

Media / SSRF Security

Host allowlisting, private IP blocking, 5MB bounded streaming, and redirect re-validation.

STATUS✓ ALL PASSING
AREA 0413 TESTS

Conversation Flow

Multi-turn state transitions, slot clarification, legacy endpoint removal, and container startup.

STATUS✓ ALL PASSING
View detailed test classes (9 test classes · 42 tests total)↓
01.MediaUrlSecurityValidatorTest
12 tests
02.EligibilityEngineTest
6 tests
03.TwilioWebhookHardeningTest
5 tests
04.AdminSchemeSecurityTest
4 tests
05.ConversationOrchestratorTest
4 tests
06.LegacyRouteEliminationTest
4 tests
07.SafeMediaDownloadServiceTest
4 tests
08.TwilioWebhookControllerTest
2 tests
09.YojnaSetuApplicationTests
1 tests

Interview Context: Passing 42 automated tests confirms that specific security regressions, SSRF vectors, and eligibility matching invariants are protected in CI. It demonstrates solid engineering hygiene rather than a theoretical guarantee against all hypothetical bugs.

08 — RELEASE STATUS & BOUNDARIES

Release Status & Known Boundaries

What is verified in V2, and what known limitations remain. The independent release gate classified Yojna Setu V2 as Release Ready With Documented Conditions.

VERIFIED CAPABILITIES7 / 7 VERIFIED
  • ✓Supabase PostgreSQL 17 live connectivity and yojna_setu schema isolation.
  • ✓82 normalized welfare schemes seeded with relational child tables.
  • ✓Groq Cloud AI connectivity for multilingual demographic extraction (gpt-oss-20b).
  • ✓Whisper Large V3 Turbo connectivity for voice-note transcription.
  • ✓42/42 automated unit, integration, and security tests passing.
  • ✓Historical plaintext credentials completely expunged from reachable repository history.
  • ✓Strict PII log masking and blind-indexed phone storage verified.
DOCUMENTED CONDITIONSEXPLICIT BOUNDARIES
Twilio WhatsApp DeliveryNOT VERIFIED IN PRODUCTION

Live production WhatsApp delivery over Twilio was not verified with live carrier traffic due to sandbox constraints. Webhook reception, HMAC verification, and response serialization are fully verified via automated suites.

Database Authorization (RLS)NOT IMPLEMENTED AT DB LAYER

PostgreSQL Row Level Security (RLS) is not currently implemented on the database tables; all access authorization is enforced strictly within the Spring Boot application service layer via authenticated database roles.

Rate Limiting ArchitecturePROCESS-LOCAL ONLY

Bucket4j rate-limiting tokens are stored in local JVM memory. While effective for single-instance deployments, horizontal scaling across multiple instances requires a centralized Redis token bucket.

09 — INTERACTIVE PORTFOLIO DEMO
PORTFOLIO DEMO · CLIENT-SIDE SIMULATORZERO CREDENTIALS OR PII TRANSMITTED

How Intake & Eligibility Work in Practice

Simulated portfolio experience — not a live government service. Select a sample persona below to walk through the 4 steps: from citizen message to extracted profile, rule checks, and matched schemes.

STEP 4 · MATCHING SCHEMES (3 QUALIFIED)FILTERED FROM 82 NORMALIZED SCHEMES

PM-KISAN Samman Nidhi

Central Welfare

Benefit: Direct income support of ₹6,000 per year in three installments for landholding farmer families.

WHY THIS CITIZEN QUALIFIES:

Active farmer, landholding criteria satisfied, income within prescribed ceiling.

Required Documents:Aadhaar CardLand Ownership Proof (Khasra/Khatauni)Bank PassbookDeadline: Rolling Annual Enrollment

Ayushman Bharat PMJAY

Central Welfare

Benefit: Health insurance cover of up to ₹5,00,000 per family per year for secondary and tertiary hospitalization.

WHY THIS CITIZEN QUALIFIES:

Low-income rural household meeting deprivation criteria.

Required Documents:Aadhaar CardRation CardIncome Certificate

UP Mahila Samarthya Yojana

State Welfare

Benefit: Financial and technical support for rural women self-help groups and agrarian enterprises.

WHY THIS CITIZEN QUALIFIES:

Female resident of Uttar Pradesh engaged in agrarian micro-enterprise.

Required Documents:Aadhaar CardUP Domicile CertificateBank Account Details
ARCHITECTURAL SUMMARY

The Engineering Takeaway

The main takeaway from Yojna Setu is separating what AI is good at from what software rules must guarantee. The LLM handles messy, multilingual conversational input, while deterministic Java code and relational database queries evaluate welfare eligibility with complete predictability.