AI/ML Engineer & Consultant

Turning complex business operations into intelligent systems

I help businesses identify high-value opportunities for AI, design practical automation strategies, and build production systems that improve operations, increase revenue, and reduce cost.

From the first process review to a working, measurable system in production.

Five peer-reviewed papers on conversational AI in sales and customer operations. Speaking at Day of Data Orlando, October 2026.

Nadia Shiroglazova

Nadia Shiroglazova
Applied AI & production systems

How the systems I build actually work
Every system on this page closes the same loop: it observes, decides, acts, measures the result, and corrects itself.
Approach
Useful, not impressive

I work at the intersection of technology and operations: understanding how a business works today, finding where intelligence or automation creates leverage, and designing systems that are practical, measurable, and connected to real workflows.

I don't build AI for the sake of having an AI feature. I focus on systems that reduce manual work, improve decisions, create qualified demand, increase operational capacity, or unlock new revenue — and I stay responsible for them through discovery, architecture, build, integration, evaluation, launch, and continuous improvement.

Background
Based
Chicago, IllinoisWorking remotely worldwide
Education
MS Data Science, NorthwesternBS Data Science, Northeastern
Previously
Boston Consulting Group · BainEnterprise data platform and AI tooling teams
Practice
100+ AI/ML engagementsProduction systems since 2023
What I do
Six places AI earns its keep
Engagements usually start with the first one and grow into the others.

AI opportunity assessment & automation strategy

  • Review sales, operations, support, and back-office processes
  • Find bottlenecks, repetitive work, and quality gaps
  • Prioritise by impact, feasibility, data readiness, and effort
  • Define target operating model, architecture, roadmap, and success metrics

Sales & agency automation

  • Lead generation and prospect discovery
  • AI-assisted and autonomous outreach
  • Voice-based qualification and follow-up
  • CRM updates, lead scoring, routing, pipeline visibility
  • Response classification and next-best-action workflows

Agentic & voice AI

  • Production voice agents for inbound and outbound work
  • Multi-step agents with tools, memory, and business context
  • Human handoff and escalation design
  • Conversation quality, latency, reliability, and cost optimisation

Predictive decision systems

  • Risk scoring and prioritisation
  • Predictive calling and capacity optimisation
  • Ranking systems and work queues
  • Operational forecasting and resource allocation
  • Interfaces that connect model output to real action

AI evaluation, quality & operations

  • LLM and agent evaluation pipelines
  • Human-in-the-loop review systems
  • Speech and operator-performance analysis
  • Monitoring, error analysis, and feedback loops

Production AI engineering

  • LLM, RAG, OCR, and vision-language systems
  • Data pipelines and model-serving workflows
  • Integrations with telephony, CRM, APIs, and internal tools
  • Self-hosted and cloud AI infrastructure
Selected work
Systems in production
Open any project to see how it is built.
10,000+
calls processed daily
~92%
accuracy on disposition validation
20% → 100%
monitoring coverage
~50%
lower monitoring headcount

Transcription, status detection, and LLM evaluation of every call against 10+ operator KPIs — script adherence, objection handling, empathy, upsell technique, resolution quality. Disputed cases go to a human arbitration loop whose decisions feed back as training signal, which lifted accuracy a further 15%. Migrating from a third-party LLM API to a self-hosted 70B model cut running costs roughly in half at the same throughput.

Speech-to-textLLM evaluationSelf-hosted 70B inferenceCelery / RedisSQL ServerHuman-in-the-loop arbitration
44%
fewer lost leads
17%
fewer dials
2–3×
sales capacity per operator

Fixed dialling settings are a bad trade: dial fast and callers wait on hold, dial slowly and operators sit idle. This controller solves the allocation on a rolling horizon instead, modelling transfer delay as an empirical lag convolution and validating through discrete-event simulation calibrated to production data. Results shown are from counterfactual replay against the previous policy. The method is published.

Receding-horizon controlInteger programmingQueueing theoryDiscrete-event simulationCounterfactual replay
~6%
partnership conversion from cold outreach
<2%
of contacts identified the agent as AI
24/7
unattended operation

Prospecting runs without a human in it: the system discovers businesses through the Places API, classifies them by category and fit, calls to qualify, then ranks by warmth from engagement signals and sends a personalised follow-up. Full call history is retained, so a prospect who calls back is resumed in context rather than re-introduced cold. A parallel pipeline discovers and negotiates influencer partnerships, selecting the commission structure per audience profile automatically.

Google Places APIAI voice qualificationLead warmth scoringResponse classificationDynamic affiliate logicCRM integration
5
horizon-specific models replacing one static one
Daily
operational re-prioritisation

A single static risk score can't tell a coordinator what to do this morning. This system runs five horizon-specific models over point-in-time features — using only what was knowable at each decision point — then surfaces the highest-risk training sequences and ranks candidate students for recovery by availability, prior disruption, training need, and current risk. Validated chronologically, with historical replay and what-if analysis built into the coordinator workflow.

Point-in-time featuresChronological validationHorizon-specific modelsRanking metricsExplainabilityCoordinator workflow
~80%
reduction in manual due-diligence effort
500+
pages per document
100%
self-hosted

Documents are classified and routed by a vision-language model with OCR fallback, then answered through a hierarchical RAG pipeline with sentence-level citation and recursive three-level summarisation. The entire stack was migrated off third-party APIs to self-hosted deployment to satisfy the client's data-residency and confidentiality requirements.

Qwen 32B VLMABBYY / Tesseract OCRHierarchical RAGSentence-level citationN8N orchestrationSelf-hosted inference
+2–3%
worldwide revenue increase
up to +50%
delivery rate on problematic addresses, depending on country
~96%
validation accuracy
15+
countries across Europe and Latin America

There is no single source of truth for addresses across 15+ countries. Each country has its own authoritative providers alongside Google Places, and each region and city carries its own formatting conventions and validity rules — so the system holds a per-country source set and a regional rule set rather than one global model.

Disagreements between sources are surfaced with plain-language reasoning instead of a bare confidence score, so a reviewer can spot-check in seconds. The most useful part is the loop: the LLM acting as judge proposes changes to the rules themselves when it sees a pattern the current rules mishandle, and a human confirms or rejects them. The rule set improves per region without anyone rewriting it by hand.

Per-country source setsNominatimGoogle Places APIRegional rule engineLLM as judgeModel-proposed rule updatesHuman review loop
Research
Published work
Peer-reviewed writing alongside the applied practice.
Published
International Journal of Advanced Artificial Intelligence Research · Vol. 3, No. 7, pp. 69–81 · July 2026 · Sole author
When AI voice bots qualify customers before handing them to human operators, dialling too fast strands callers on hold and dialling too slowly leaves operators idle. The paper replaces fixed dialling settings with a receding-horizon integer-programming controller that models transfer delay as an empirical lag convolution — capturing most of the throughput of aggressive dialling while cutting abandonment and hold times. It is the method behind the predictive calling system above.
Receding-horizon controlInteger programmingQueueingProduction-calibrated simulation
Read the paper · DOI 10.55640/ijaair-v03i07-07
Earlier publications
2025
Designing Conversational AI Scripts for Emotional Engagement and Customer Readiness Detection in Sales Calls
2025
Balancing Automation and Human Intervention in Early-Stage Sales Calls
2024
Trustworthy AI for Conversational Customer Service Systems
2024
Automated Evaluation of Sales Calls Using Conversational AI

These four papers are being reissued with a new publisher over the next one to two months. Direct links will be added here as each becomes available again — happy to send copies in the meantime.

Research projects
2025
LLM knowledge gaps in low-resource programming languages
Supervised faculty research, Northeastern University
Why models that write reliable Python break down on Julia, and which intervention actually closes the gap — domain fine-tuning, RAG over language documentation, or in-context library injection.
2024
Adversarial attacks on code-generation models
Northeastern University
Black-box attacks against StarCoder and Code LLaMA using prompt injection and crafted prefixes, evaluating how fill-in-the-middle completion can be steered toward vulnerable code.
How I work
Six stages, in order
1

Discover

Learn how the business actually runs today — the process, the people, the data, and where the friction really sits.

2

Design

Choose the opportunity worth pursuing, define success metrics, and design the system and the operating model around it.

3

Build

Develop models, agents, and pipelines against real data, with evaluation in place from the first version.

4

Integrate

Connect the system to the tools people already use — telephony, CRM, databases, internal interfaces.

5

Measure

Compare against the baseline honestly: backtesting, counterfactual replay, human review, and cost per outcome.

6

Improve

Feed results back in. Recalibrate, extend coverage, and remove the failure modes that surface in production.

Technical depth
Under the hood
For the engineers in the room.
LLMs & AI agents
Tool use & function callingMulti-agent workflowsConversational systemsPrompt & model orchestrationMemory and business context
RAG & document intelligence
Retrieval & embeddingsVector searchOCRVision-language modelsHierarchical summarisationCitation grounding
Predictive modelling
ClassificationRankingForecastingRisk scoringCapacity optimisationQueueing models
Voice AI & telephony
Real-time conversationCall routingHuman handoffCampaign pacingLatency optimisationCost per contact
Data & integrations
PythonSQLAPIsCRM systemsTelephony platformsWorkflow orchestrationDashboards
Evaluation & monitoring
BacktestingHuman-in-the-loop reviewError analysisCalibrationQuality monitoringFeedback loops
Infrastructure
Self-hosted inferenceCloud AI servicesDockerDistributed workflowsProduction deployment
Speaking
On stage
Available for conferences, panels, and internal briefings.
Speaking next
Beyond Model Accuracy: Designing a Risk-Aware LLM Agent with Confidence Gating and Human Escalation
Day of Data Orlando 2026 · 17 October 2026 · Seminole State College, Sanford, Florida
Other topics I speak on
Building AI agents that work in production
From AI demo to business-critical system
Predictive calling and the future of AI-assisted sales
Automating customer operations with voice AI
Evaluating AI systems in real business workflows
Using AI to understand and improve operator performance
Address validation and human-in-the-loop AI
Practical AI strategy for sales and operations teams

Build something useful

Have a complex process, an underused data source, or an AI opportunity you're not sure how to approach? Let's discuss what could create real value.

Messages go straight to [email protected]. Prefer email? Write directly.