Capabilities
Full-stack AI, not just a model
A model is roughly 20% of a working AI system. The other 80% — ingestion, retrieval, evaluation, serving, monitoring, integration — is what decides whether it survives contact with production. We build all of it.
Capability 01
Generative AI & LLMs
Most "generative AI" projects fail for an unglamorous reason: nobody defined what correct output looks like before shipping. We start from the evaluation set. Domain experts produce a gold standard for the task — drafting, extraction, classification, summarisation — and that set becomes the contract the system is held to, not a vibe check in a review meeting.
From there the engineering question is how much model you actually need. A well-constructed prompt program with structured output and retrieval beats a fine-tune more often than vendors admit, and it is dramatically cheaper to maintain. When a fine-tune is warranted — domain vocabulary, output format rigidity, latency or cost pressure at volume — we use LoRA or QLoRA against an open-weight base so the artefact is yours and runs on hardware you control.
Everything ships with constrained decoding against an explicit schema. A language model that can emit free-form text into a downstream system is an outage waiting for a trigger; one whose output is validated against a Pydantic model before it leaves the service is a component you can build on.
Technology depth
| Dimension | What we work with |
|---|---|
| Model selection | Hosted (Claude, GPT-class) and open-weight (Llama 3.3, Mistral, Qwen) benchmarked on your task, not on public leaderboards |
| Adaptation | Prompt programs (DSPy), few-shot optimisation, LoRA / QLoRA fine-tuning, continued pre-training for specialist vocabularies |
| Output control | JSON-schema constrained decoding, grammar-constrained generation, Pydantic validation, reject-and-retry loops |
| Safety | Input/output classifiers, PII redaction, jailbreak red-teaming, per-tenant policy enforcement |
| Serving | vLLM, TensorRT-LLM, quantisation (AWQ / GPTQ / INT8), speculative decoding, request batching |
| Evaluation | Gold-set accuracy, LLM-judge with human calibration, regression gates in CI, per-segment reporting |
Capability 02
Retrieval & Knowledge Systems
Retrieval is where the majority of enterprise AI value sits, and where the majority of prototypes quietly fail. The gap is almost never the embedding model. It is ingestion: tables flattened into unusable prose, multi-column PDFs read in the wrong order, OCR failing silently on scanned pages and producing plausible garbage that gets indexed and cited.
We build ingestion as the first-class component it is. Layout-aware parsing preserves table structure and reading order. OCR runs behind a confidence gate that routes low-confidence pages to human review rather than degrading the index silently. Chunking follows document structure — a clause, a table, a policy section — with contextual headers prepended so each chunk is interpretable on its own.
Retrieval itself is hybrid by default. Dense vectors catch paraphrase; BM25 catches the exact entity names, account numbers and defined terms that dense retrieval reliably confuses. Fused and cross-encoder reranked, then generated under enforced citation: every claim resolves to a document, page and region a human can verify in one click, or the response is rejected by the validator and escalated. Abstention is a first-class, calibrated output — a system that cannot say "not found in this file" is not deployable in a regulated workflow.
Technology depth
| Dimension | What we work with |
|---|---|
| Ingestion | Layout-aware parsing, table-structure preservation, confidence-gated OCR (PaddleOCR / Textract), deduplication, incremental CDC indexing |
| Chunking | Structure-aware boundaries, contextual header injection, small-to-large parent expansion, multi-granularity indexing |
| Retrieval | pgvector / Qdrant dense search, OpenSearch BM25, reciprocal-rank fusion, metadata and permission pre-filtering |
| Reranking | Cross-encoder reranking (Cohere Rerank, bge-reranker), retrieve-50-rerank-to-8 topologies, score-based abstention thresholds |
| Grounding | Schema-enforced citations validated against the retrieved set, page-region resolution, immutable audit logging |
| Evaluation | Recall@k, MRR, nDCG for retrieval; faithfulness, answer correctness, citation accuracy and false-abstention rate for generation (Ragas + custom harness) |
Capability 03
Computer Vision
In industrial vision, the highest-leverage engineering decision is usually lighting, not architecture. A defect class that is nearly invisible under diffuse light can become trivially separable under strobed darkfield at the right angle — which is why we specify the imaging rig alongside the model and work directly with your controls engineers in the architecture phase.
Models are selected against the latency budget rather than the benchmark table. A line running 240 parts per minute with a reject gate 1.2 metres downstream gives you a hard budget from trigger to PLC decision, including acquisition and round-trip. That budget determines quantisation strategy, edge hardware and whether a two-stage detect-then-classify pipeline is affordable — and it is usually worth paying for, because per-region explanations are what make quality engineers trust the system.
Annotation is the hidden cost, so we attack it directly: SAM 2-assisted labelling where inspectors correct candidate masks instead of drawing them, which has cut our per-part annotation time by roughly 75% in practice. And we monitor input distributions, not just predictions — image-statistic drift per exposure channel catches a degrading strobe LED weeks before it shows up as an accuracy regression.
Technology depth
| Dimension | What we work with |
|---|---|
| Acquisition | Camera and optics specification, strobed / darkfield / multi-angle lighting design, hardware triggering, GenICam integration |
| Tasks | Defect detection and classification, segmentation, dimensional measurement, OCR/OCV, anomaly detection for unseen defect classes |
| Architectures | YOLO family, RT-DETR, SAM 2, EfficientNet / ConvNeXt classifier heads, PatchCore and autoencoder anomaly models |
| Annotation | SAM 2-assisted labelling, active learning loops, inter-annotator agreement measurement, synthetic defect augmentation |
| Edge serving | TensorRT INT8/FP16 on Jetson Orin, ONNX Runtime, OpenVINO on industrial PCs, fleet updates via Balena |
| Integration | OPC UA and Modbus to PLCs, MQTT event streams, reject-gate actuation, operator HMI with per-region explanations |
Capability 04
Predictive Analytics
Forecasting in real operations is rarely a time-series textbook problem. It is a hierarchy problem, an intermittency problem and a promotions problem simultaneously — and the three interact. A model that nails chain-level demand and cannot be reconciled to store level is unusable, because finance plans against one number and execution happens against the other.
So we forecast distributions, not points, and reconcile by construction. Quantile objectives give replenishment policy the service-level quantile it actually needs — essential for the long intermittent tail where a mean forecast is actively misleading. MinT reconciliation makes every level of the hierarchy sum, which ends a surprising amount of organisational argument. Promotions are explicit features owned by a calendar, never noise absorbed into a trailing average.
Gradient-boosted trees remain the right default for tabular forecasting at this scale — they outperform deep sequence models on most real retail and operational data, train in minutes rather than hours, and expose feature contributions that a planner can interrogate. That last property is what drives adoption: a forecast that shows its top drivers and last year's actual gets trusted; an unexplained number gets overridden into irrelevance.
Technology depth
| Dimension | What we work with |
|---|---|
| Problem classes | Hierarchical demand forecasting, intermittent demand, capacity and workforce planning, churn and credit risk scoring, failure prediction, price and route optimisation |
| Models | LightGBM / XGBoost with quantile objectives, Prophet and statsmodels baselines, Croston / TSB for intermittency, survival models, Bayesian structural time series |
| Hierarchy | Bottom-up forecasting with MinT / OLS reconciliation, store-cluster segmentation, cold-start via attribute similarity |
| Features | Multi-window lags, promotional calendar and depth, relative price, holiday proximity, weather and local-event feeds, feature stores for train/serve parity |
| Validation | Rolling-origin backtesting, WAPE / MASE / pinball loss, segmented accuracy by velocity class and promotional status, post-override benchmarking |
| Delivery | Ray distributed training, Airflow / Dagster orchestration, Great Expectations input gates, planner UI with explanations and reason-coded overrides |
Capability 05
AI Agents & Automation
Agents are the capability most oversold and least often deployed successfully, and the reason is scope. An agent given broad autonomy over real systems fails in ways that are hard to debug and expensive to undo. An agent given a bounded task, a typed tool surface and an approval gate on consequential actions is just reliable software with a flexible front end.
We build the bounded kind. Tools are explicit, typed and individually permissioned. State lives in a durable workflow engine — Temporal, usually — rather than in a conversation history, so a step that fails on a transient API error resumes rather than restarting, and every run is replayable after the fact. Anything irreversible (sending money, emailing a customer, writing to a system of record) sits behind an approval gate until the measured success rate justifies removing it.
Observability is the deciding factor between an agent you can operate and one you quietly turn off. Every tool call, argument, result, retry and decision is traced through OpenTelemetry and inspectable per run. When an agent does something surprising in production — and it will — you need the trace, not a transcript.
Technology depth
| Dimension | What we work with |
|---|---|
| Patterns | Single-task bounded agents, supervisor/worker decomposition, plan-then-execute, human-in-the-loop escalation, batch reconciliation agents |
| Tooling | Model Context Protocol (MCP) servers, typed tool schemas, per-tool permissioning, idempotency keys, dry-run modes |
| Orchestration | LangGraph for control flow, Temporal for durable execution and retries, Redis for state, Postgres for run history |
| Guard-rails | Approval gates on irreversible actions, spend and step budgets, loop detection, allow-listed targets, rollback procedures |
| Observability | OpenTelemetry tracing of every tool call, per-run replay, cost and latency attribution, success-rate dashboards by task type |
| Integration | ERP, CRM, ticketing and data-warehouse connectors; webhook and queue-driven triggers; SSO and row-level permission inheritance |
Capability 06
MLOps & Deployment
This is the half of the work that determines whether a model is still running in eighteen months. It is also the half most consultancies leave for the client to figure out, which is how organisations end up with a Jupyter notebook in production and one person who knows how to restart it.
We ship reproducible pipelines: versioned data, versioned code, versioned models, versioned prompts, with a registry that can tell you exactly what was serving on any given date. Evaluation runs in CI as a gate — a change that regresses a tracked metric beyond tolerance does not merge. That single control prevents more production incidents than any amount of monitoring, because it catches the regression before it ships.
Deployment topology follows your constraints rather than our preferences. Roughly half our deployments run inside a customer VPC; we regularly deploy fully on-premises, including air-gapped facilities with no outbound network access. That constraint is identified in week two of the engagement, because it determines model selection — and discovering it in week eight is how AI projects die.
Then we hand it over properly: runbooks, dashboards, alert thresholds, an on-call escalation path, and training for the team that owns it. The measure of a good engagement is that you do not need us afterwards.
Technology depth
| Dimension | What we work with |
|---|---|
| Reproducibility | DVC / LakeFS data versioning, MLflow experiment tracking and model registry, prompt and config versioning, pinned container builds |
| CI/CD | GitHub Actions pipelines, evaluation gates on pull requests, canary and blue/green model rollout, tested rollback paths |
| Serving | Kubernetes with KServe or bespoke FastAPI services, vLLM for LLM inference, autoscaling, GPU scheduling and bin-packing |
| Monitoring | Input and prediction drift detection, accuracy tracking against delayed labels, latency / cost / token dashboards in Grafana, alerting with runbook links |
| Deployment targets | Customer VPC (AWS / Azure / GCP), on-premises Kubernetes, edge devices, fully air-gapped environments with offline model distribution |
| Handover | Infrastructure-as-code (Terraform), runbooks, architecture decision records, operator training, optional advisory retainer |
Not sure which of these you need?
Most engagements start with a problem, not a capability. Describe the process that is costing you the most and we will tell you which of these — if any — is the right tool.