India  ·  Est. 2012  ·  Delivering to US, UK & Asia-Pacific
+91 93635 07420 info@maalkum.com

AI Data &
Technology.

We handle the full spectrum of AI data operations — from annotation and evaluation to building the software systems that run on top of that data. One team covering both sides.

Get in Touch →
AI Data Operations

Annotation, evaluation, and validation
for AI and LLM systems.

We generate, annotate, validate, and evaluate the training data that AI systems depend on. Schema-compliant delivery with documented fix cycles — not crowdsourced annotation.

We annotate multi-turn agentic conversation trajectories for AI safety evaluation — tool call labelling, PII entity identification, and cumulative scoring across turns.

  • Multi-turn trajectory annotation per client-defined schema
  • CF and PA scoring across turns
  • PII entity identification — 31+ entity types
  • Tool IO label validation
  • Schema compliance tracking and documented fix cycles

What clients send us

Schema/SOPYour annotation spec — we read it fully before touching any task.
Raw trajectoriesAgentic conversation data in JSON or your format.
Pilot batchSmall pilot first, QA feedback, fix, then scale.

We review, rank, and correct AI outputs — building the preference datasets that make models safer and better aligned with human values.

  • Response ranking and preference pair creation
  • AI output evaluation against defined rubrics
  • Safety and policy compliance assessment
  • Domain-specific evaluation — coding, legal, medical, general
  • Iterative feedback loop support

Suitable for

LLM safety teamsBuilding preference datasets for safety alignment.
Model evaluation teamsSystematic human evaluation at scale.
Fine-tuning pipelinesHigh-quality ranked pairs for preference fine-tuning.

We prepare SFT datasets from scratch or clean and validate existing datasets — following your format spec and quality rubric precisely.

  • Instruction-response pair creation per your schema
  • Domain-specific data — coding, legal, medical, general
  • Format validation and schema compliance checks
  • Deduplication, quality filtering, and diversity checks
  • Cleaning and repair of existing SFT datasets

Suitable for

LLM fine-tuning teamsDomain-specific instruction datasets for supervised fine-tuning.
Model customizationAdapting base models to specific domains or task types.
Dataset repairCleaning and fixing existing datasets before training.

We generate synthetic prompt-response pairs for LLM training and validate them before they enter your pipeline — catching fabrication, schema errors, and SOP violations systematically.

  • Prompt-response pair generation per client SOP
  • Bulk QA validation with structured reports
  • Fabrication detection — hallucinated fields, invented thresholds
  • Severity classification — hard violations vs warnings
  • Custom validator tooling built per engagement

What makes this different

We build toolingBulk QA validator built as part of the engagement.
We do bothGeneration and QA in one team — no handoff errors.
We iterate fastValidator logic updated in real time.

We provide structured human evaluation of LLM outputs — scoring responses against defined rubrics, running model comparisons, and executing benchmark tasks.

  • Human evaluation against scoring rubrics
  • Side-by-side model comparison and preference scoring
  • Benchmark task execution and result documentation
  • Agentic workflow evaluation — tool use, multi-step reasoning
  • Structured eval reports with inter-rater reliability metrics

Suitable for

Model evaluation teamsHuman evals alongside automated benchmarks.
Research labsStructured evaluation for model papers and capability assessments.
AI product teamsPre-launch human evaluation of model-powered features.

We provide structured human review at every stage where automation alone is not sufficient — from pre-training data validation to live output monitoring.

  • Pre-training data review and quality gates
  • Live output sampling and quality monitoring
  • Edge case escalation and resolution workflows
  • Feedback loop management between model and human reviewers
  • Ongoing HITL retainer — scale with your pipeline

Why HITL matters

Automation has limitsEdge cases and policy decisions require human judgment.
Continuous improvementHITL feedback loops make models better over time.
Audit trailDocumented human review for regulated industries.

We review and validate existing data, outputs, and workflows against defined standards — prompt QA, schema validation, SOP compliance, tool-calling validation, and red teaming.

  • Prompt QA validation against quality and format standards
  • SOP compliance review — outputs checked against your SOP
  • Schema validation — data structures checked against required schemas
  • Multi-turn conversation QA across dialogue turns
  • Tool-calling workflow validation
  • Red teaming — adversarial testing for safety issues
  • AI safety labeling for policy compliance evaluation

Why this is separate

Different buyerAI safety teams need validation — separate from annotation procurement.
Higher stakesSchema errors in production cost more than the QA to catch them.
Benchmark supportWe build and execute benchmark evaluation environments.

We collect and annotate data for AI training — gathering voice recordings, images, and text data, then labelling it for machine learning pipelines.

  • Voice and audio data collection for speech recognition
  • Image and photo data collection for computer vision
  • Image classification, object detection, segmentation
  • Text classification, NER, sentiment labelling
  • Video frame annotation and action labelling
  • PII identification and redaction
  • Multilingual annotation support

Quality controls

Inter-annotator agreementMultiple annotators on sensitive tasks.
QA samplingRandom sample checks on every batch before delivery.
Fix cyclesDocumented rework on client-flagged issues.

We provide structured review of AI-generated outputs for quality, safety, bias, and policy compliance — systematically, not just spot-check.

  • AI output quality evaluation against rubrics
  • Safety and harmful content assessment
  • Bias detection and flagging
  • Policy compliance review
  • Multimodal review — text, image, combined outputs
  • Structured issue reports with severity classification

Suitable for

AI product teamsPre-launch content review for generative AI features.
Trust & Safety teamsOngoing moderation for AI-generated content at scale.
Model evaluationSystematic human evaluation alongside automated evals.

AI Engineering

Production AI systems —
not demos or prototypes.

We build RAG pipelines, agentic systems, LLM fine-tuning pipelines, and GPU model serving infrastructure — shipped and running in production.

Retrieval-Augmented Generation systems built for accuracy and reliability — multi-source retrieval, reranked, and monitored.

  • Vector database integration — Milvus, ChromaDB, Pinecone
  • Parallel async retrieval across multiple sources
  • Query rewriting for better recall
  • Reranking pipelines for retrieval precision
  • Circuit breakers, retry logic, and caching for reliability
  • Retrieval quality evaluation and optimization

Production RAG

Multi-sourceParallel retrieval across vector databases and live external sources.
Quality optimizationQuery rewriting + reranking reduces hallucination systematically.
ReliabilityCircuit breakers ensure the pipeline holds under real load.

Multi-agent systems with stateful reasoning, tool use, and autonomous task execution using LangGraph and custom state machines.

  • Multi-agent systems with specialized role separation
  • Stateful reasoning with persistent memory across steps
  • Tool use — web search, APIs, database queries, code execution
  • State machine architecture for deterministic routing
  • Intent classification and dynamic agent routing
  • Full observability — tracing, metrics, production health checks

What we have built

Multi-agent RAGOrchestrated specialized agents via state machine architecture.
AI sales automation3-stage pipeline: company analysis → strategy → content generation.
Enterprise productivityMicrosoft 365 integration with LLM-driven prioritization.

We build and run fine-tuning pipelines for domain adaptation and task specialization using LoRA/PEFT — full lifecycle from dataset generation to evaluation.

  • LoRA and PEFT fine-tuning for domain adaptation
  • Automated training dataset generation pipelines
  • Contrastive learning for embedding model improvement
  • Evaluation benchmarks with measurable quality metrics
  • Experiment tracking with W&B and DVC
  • Model versioning and performance comparison

When fine-tuning makes sense

Domain specializationWhen a base model needs to learn your terminology or reasoning patterns.
Embedding improvementWhen retrieval quality needs to improve beyond prompt engineering.
Cost reductionSmaller fine-tuned models can outperform larger general models on specific tasks.

GPU-optimized model serving for production workloads — multiple large models concurrently with efficient memory allocation and full monitoring.

  • High-throughput inference with vLLM and SGLang
  • Multi-model concurrent serving with memory optimization
  • Embedding models and rerankers alongside LLMs on shared GPU
  • Production monitoring — Prometheus, Grafana, alerting
  • Model evaluation, versioning, and migration across generations

Infrastructure

GPU servingLarge models alongside embedding models on optimized hardware.
Full observabilityPrometheus metrics, Grafana dashboards, alerting across all endpoints.
Managed migrationsModel updates with minimal service interruption.

Custom conversational agents built for production — with memory, routing, fallback handling, and full observability. Not a thin API wrapper.

  • Multi-turn conversation memory and context management
  • Domain-specific knowledge base integration
  • Multi-provider LLM support — OpenAI, Anthropic, open-source
  • Fallback handling, intent routing, and escalation flows
  • Production monitoring and conversation analytics

Built for production

Not a wrapperCustom architecture with memory, routing, fallback, and observability.
Multi-providerNot locked to one LLM provider — swap models without rebuilding.
MonitoredTracing, logging, and alerting built in from day one.

AI-powered automation flows connecting your systems — n8n, CRM integrations, webhook pipelines, document intelligence, and LLM-powered generation.

  • n8n-based workflow orchestration
  • CRM and ERP integrations with AI-driven logic
  • Multi-tier PDF extraction with fallback chains — handles corrupted and scanned documents
  • Bulk document ingestion with parallel processing
  • Microsoft 365 integrations — Outlook, Teams, Calendar
  • Intelligent email and document generation systems

What we automate

Document workflowsIntelligent document generation and routing triggered by business events.
Sales prospectingLLM-powered research, strategy, and personalised outreach — end to end.
Enterprise productivityMicrosoft 365 integrations that surface what needs attention.

Technology Services

Software that ships —
web, mobile, cloud.

We build web applications, mobile apps, and cloud infrastructure. Dedicated engineering teams or fixed-scope projects.

We build web applications from the ground up — from simple sites to complex platforms with authentication, data pipelines, and real-time features.

  • Frontend: React, Next.js, Vue, TypeScript
  • Backend: Python (FastAPI/Django/Flask), Node.js, Java, .NET
  • Database design — PostgreSQL, MySQL, MongoDB, Redis
  • RESTful and GraphQL API development
  • Authentication, authorization, and security hardening
  • CMS integration — headless and traditional

Examples

SaaS platformsMulti-tenant apps with billing, dashboards, and role-based access.
Data dashboardsReal-time analytics and reporting tools.
Custom portalsClient and vendor portals with workflow automation.

We build cross-platform mobile apps using Flutter and React Native — one codebase, both iOS and Android.

  • Flutter — iOS and Android from a single codebase
  • React Native — JavaScript-based cross-platform apps
  • Offline-first architecture where required
  • Push notifications, location, camera, biometric integrations
  • App Store and Google Play deployment support
  • Ongoing maintenance and update cycles

Our approach

Cross-platform firstFlutter and React Native deliver iOS + Android at a fraction of the cost of two native apps.
API-connectedAll apps connect to your backend — existing or built by us.
Design includedUI/UX design is part of the build.

We design, build, and manage cloud infrastructure on AWS, GCP, and Azure — from your first deployment pipeline to a scalable production environment.

  • AWS, GCP, and Azure infrastructure setup and management
  • Docker containerization and Kubernetes orchestration
  • CI/CD pipeline setup — GitHub Actions, GitLab CI, Jenkins
  • Infrastructure as Code — Terraform, Pulumi
  • Monitoring, logging, and alerting — Datadog, Grafana, CloudWatch
  • Security hardening and cost optimization

Cloud platforms

AWSEC2, ECS, Lambda, RDS, S3, CloudFront, SageMaker.
GCPGKE, Cloud Run, BigQuery, Vertex AI, Cloud Storage.
AzureAKS, Azure Functions, Cosmos DB, Azure ML, Blob Storage.

We build automated test suites that catch issues before your users do — integrated into your CI/CD pipeline so every deployment is validated.

  • Unit, integration, and end-to-end test automation
  • Selenium, Playwright, Cypress for web testing
  • API testing with Postman, Newman, RestAssured
  • Performance and load testing
  • Mobile test automation — Appium, Detox
  • Manual QA for complex user flows and edge cases

Our approach

Shift leftTesting starts early — not after the build is done.
Automated firstManual testing for what cannot be automated. Automation for everything repeatable.
CI/CD integratedEvery test runs on every push. No surprises at deployment.
What We Have Delivered

Production work across all three areas.

AI Data Operations

Agentic annotation, RLHF, SFT data, synthetic prompt generation and QA, LLM evaluation, prompt QA validation, schema validation, and HITL operations for clients building AI and LLM systems.

AI Engineering

Production RAG pipelines, multi-agent systems built with LangGraph, LLM fine-tuning pipelines, GPU model serving with vLLM and SGLang, document intelligence platforms — shipped and maintained in production.

Technology Services

Full-stack web applications including billing systems, job boards, and data processing tools. Mobile apps. Cloud infrastructure. AI/ML engineering integrated into client products.

Ready to discuss your project?

Send us a brief — we respond within 24 hours.

Get in Touch →