Where does inference actually run?
Inside your perimeter. Your data centre, your own cloud tenancy in-region, or isolated in-country hardware. There is no BizfyLabs-operated endpoint in the path, because none was built. Nothing calls home.
FAQ
Inference location, air-gap, training, licences, residency versus sovereignty, keys, pricing, hardware and how a paid proof of value actually runs.

Inside your perimeter. Your data centre, your own cloud tenancy in-region, or isolated in-country hardware. There is no BizfyLabs-operated endpoint in the path, because none was built. Nothing calls home.
Yes, and that is the base case rather than a hardened option. Model weights, licence text and documentation ship as files on disk. Updates arrive as signed offline bundles on your schedule. Diagnostics produce a file you inspect before it goes anywhere.
Not by default. Your data stays inside your environment. Depending on the use case, we can use retrieval, prompting, configuration, fine-tuning or model training, without sending your data to a third-party model provider.
Only Apache 2.0 and MIT weights, vision and OCR, LLM, embeddings and reranking. Every component, version and licence is published, and the licence text ships on disk. An SBOM covering models as well as code is generated per release. Nothing carries a usage gate or a policy a vendor can revise after you deploy.
Residency answers where the bytes sit. It does not answer who controls the model reading them. If the weights are operated by a vendor in another jurisdiction, you have residency without sovereignty, and under CBUAE and the Health ICT Law, that gap is yours to explain, not the vendor's.
You do, in every tier that touches your data. In managed single-tenant deployments the infrastructure sits in our account contractually assigned to you, and the keys remain yours.
Fixed annual licence, per environment for the platform, tiered by volume for application modules. No per-token pricing, no per-page metering, no usage caps. Deployment is a fixed-fee package, never quoted in man-days. Your cost does not move when your volume does.
Starter configuration is 8 vCPU and 32 GB with no GPU. A GPU purchase is never a precondition for beginning. Sizing tables by hardware tier are published with the test set and date they were measured on.
You don't, from a marketing page, which is why we don't publish undated numbers. Every figure we release carries its test set, composition, date, model version and scoring method. The faster answer is the paid proof of value: 30 to 45 days on your own documents, in your own environment, with success criteria agreed in writing before we start.
We start with the business problem and constraints, then design the model, data, application and deployment architecture. We build and evaluate the system against representative workloads before deploying it inside your environment. Existing systems can be connected through APIs, databases and enterprise integrations.
In your office, on your documents, with the cable pulled out. No competitor's sales engineer can do this.