Pratik Pawar — AI & Backend Engineer
B.Tech AI & Data Science at VIT Pune, building document intelligence and retrieval systems.
I build systems that read messy documents and make them searchable — extraction, normalisation and semantic tagging over sources that were never meant to be parsed. Most of that work is Python, with the retrieval side sitting on embeddings and vector search.
Currently a Google Summer of Code 2026 contributor with HumanAI, where the funding-intelligence pipeline lives. Backend work is Flask and Express behind REST APIs, and I care about the pytest suite and the CI as much as the model.
Sections
- Workstation — Experience · Projects · CV
- Television — Games · Profile
- Bookshelf — Skills · Tech Stack
- Record Player — Now Playing
- Guitar — Covers · Recordings
- Contact — Send a message
Experience
- Google Summer of Code 2026 Contributor, HumanAI (May 2026 — Present) — End-to-end document intelligence pipeline that extracts, normalises and semantically tags research funding opportunities from heterogeneous web and PDF sources. Architected and delivered the full extraction, normalization and semantic tagging pipeline across heterogeneous web and PDF sources Increased funding discovery coverage by 70%, and lifted semantic tagging from 0.62 to 0.85 F1 using ontology-based matching and all-mpnet-base-v2 embeddings Built a pytest suite — unit tests per module plus end-to-end integration tests, with fixtures and parametrized cases covering varied document formats and malformed inputs Wired GitHub Actions CI to run the suite on every push, and instrumented ingestion with structured logging to isolate parser failures on non-conforming sources Implemented retrieval, ranking and metadata enrichment stages to improve relevance of results for researchers
Projects
- Agentic AI RFP Responder (2026) — Modular multi-agent system that automates Request for Proposal analysis and response generation, cutting preparation from weeks of manual review to days. Built with Python, RAG, Vector Search, LLMs.
- Hospital Readmission Prediction System (2025) — 30-day readmission risk service trained on 101,766 patient records, deployed behind a REST API and paired with generated clinical recommendations. Built with Flask, Express.js, REST APIs, Scikit-Learn, Gemini 2.5.
- Legal Document Intelligence & Welfare Platform (2025) — OCR and information-extraction pipeline over unstructured legal records, feeding an explainable decision support system for welfare scheme recommendation. Built with TrOCR, GPT-4, LLMWhisperer, Scikit-Learn, Explainable AI.
Education
- B.Tech, Artificial Intelligence and Data Science — VIT Pune (VIIT) (2027), 9.21 CPI
- Class XII — Jadhavar Junior College, Pune (2023), 74%
- Class X — Muktangan English School (2021), 95.4%
Skills
- Languages: Python, C++, SQL
- AI / ML: PyTorch, TensorFlow, Scikit-Learn, OpenCV, Hugging Face, Deep Learning, NLP, LLMs, RAG, Semantic Search, FAISS, Pinecone, Fine-Tuning, Computer Vision, OCR, Explainable AI
- Backend & Databases: Flask, Express.js, REST APIs, PostgreSQL, MySQL, MongoDB, Firebase
- Testing & Automation: pytest, GitHub Actions CI, Structured logging, Git, Linux
- Cloud & Infrastructure: AWS, GCP, Docker, Kubernetes
- CS Fundamentals: Data Structures & Algorithms, OOP, DBMS, Operating Systems, Computer Networks, Artificial Intelligence
Contact
@PratikPawar1401