Inceptron
Founded 2021 · 20 employees on LinkedIn · 6 known investors
Inceptron provides a platform for deploying, optimizing, and scaling AI models in production, offering compiler-driven performance optimization and inference infrastructure for machine learning teams.
Also known as Inceptron AB
Founders & leadership
Inceptron was founded in 2021 by Steffen Malkowsky.

Investors · 6
Also in the syndicate · 1
Company profile
researched Aug 2026Inceptron operates a platform for running AI models in production, combining serverless inference endpoints, dedicated GPU deployments and a proprietary optimization compiler. Users can call open-source or fine-tuned models through an OpenAI-compatible API at api.inceptron.io, import custom checkpoints, or start from a curated model library, then create versioned endpoints with API keys, access controls and rollout policies. The hosted catalogue includes large open-weight models such as GLM-5.2, Kimi-K2.7 Code, Kimi-K2.6, DeepSeek-V4 Flash and MiniMax-M2.5, listed with per-token pricing, quantization format, parameter size, context length and throughput figures, alongside Llama and Qwen family models.
The optimization layer is built around a compiler that fuses computation graphs, auto-tunes kernels and plans memory for a target hardware configuration, plus performance-aware compression through quantization and pruning. Company material describes automated passes including shift-based reparameterization, mixed-precision optimization, memory and cache tuning, and workload-specific search, applied across GPUs, CPUs and accelerators such as AWS Inferentia and FPGAs, for training, diffusion-model deployment and edge or cloud inference. Operational features include dynamic request batching, autoscaling, scheduled inference, integrated logging and tracing, and connections to cloud storage buckets, CI/CD and external observability stacks.
The platform is positioned for enterprise use with team controls, encryption and hardened isolation, ISO 27001 and GDPR compliance claims, data-residency controls including EU-only residency, SLA-backed uptime and zero-retention data handling.
Founding story
The founding team came from machine-learning research backgrounds, with experience described as spanning AI labs and infrastructure companies, and built an optimization compiler aimed at making model deployment across varied hardware and frameworks more efficient [5][7]. Named co-founders include Steffen Malkowsky, listed as CTO and founder, and a co-founder identified as Anton M. [5][7].
Business model
Inceptron sells access to its inference platform directly to engineering and machine-learning teams, with self-serve serverless usage alongside contracted dedicated deployments and enterprise add-ons such as custom regions, private networking and premium support quoted separately [0][4].
Serverless inference is billed per token, with separate input and output token rates rounded to the nearest 1,000 tokens per request batch (for example Llama-3.3-70B Instruct at $0.10 per million input and $0.30 per million output tokens); dedicated deployments are billed hourly per GPU pro-rated to the minute (for example an H100 80 GB instance at $3/hour). Commitment discounts are offered at 5% for one month, 10% for six months and 20% for twelve months, applying to both dedicated and serverless usage, with overage charged at on-demand rates and unused prepaid amounts rolling over within the commitment window [4]. Published catalogue prices range from $0.13 per million input tokens for DeepSeek-V4 Flash to $1.20 for GLM-5.2 [3].
Traction
Reported headcount is in the 11-20 range, with employees across engineering, design and sales functions [7]. The compiler is in an early-access phase and the hosted platform serves a catalogue of open-weight models with published pricing and throughput figures [0][3].
Latest developments
The compiler has been opened for early access with automated model compilation, and new models including DeepSeek V4 Flash have been added to the serving catalogue [0][3].
▸Full profile — market position, technology, go-to-market, geography, history, risks & controversies
Market position
Positions itself on price-performance for AI inference, offering hosted open-weight models at published per-token rates and pitching compiler-driven optimization as the source of lower latency and cost relative to unoptimized serving [0][3].
Company material emphasises a compiler that is applied to models without usage restrictions or added overhead, automated benchmarking and compile-time profiling, accuracy-preserving or use-case-tuned optimization delivered as a modular system, and engineering support alongside the hosted platform [0][7].
Technology
A proprietary inference-optimization compiler performs graph-level fusion, hardware-aware compilation, kernel auto-tuning and memory planning, combined with model compression via quantization and pruning. Auto-tuning uses ML agents together with Bayesian optimization to search for efficient algorithm implementations for a given model and target hardware, with tuning results aggregated in databases and as model weights to improve subsequent tuning. The serving layer adds dynamic batching, autoscaling, scheduled inference and unified observability correlating logs, metrics and traces across functions, containers and workloads, and exposes an OpenAI-compatible chat completions API [0][2]. Optimization targets include GPUs, CPUs and accelerators such as AWS Inferentia and FPGAs, with support for hybrid clouds and VPCs [7].
Go-to-market
Self-serve sign-up and API access on the website, a documentation site with quick-start code samples, a public model catalogue with a playground, an early-access program for the compiler, a Discord community for support and feature requests, and a sales contact form routing enquiries by use case (custom deployment, serverless inference, model/hardware optimization, fine-tuning) [0][2][3][4].
Machine-learning and product engineering teams deploying open-source or fine-tuned models in production, spanning prototyping and workload evaluation through to enterprise buyers requiring SLA-backed uptime, EU data residency and custom deployments [0][3][4].
Geography
Headquartered in Sweden, at Scheelevägen 15, Alpha 3 (Ideon Agora), Lund, Skåne County [5][7]. The platform offers region selection and EU-only data residency for compute and logs [3][4].
History
Founded in 2021 in Sweden [5][7]. The company announced a pre-seed round in November 2023 while developing an AI optimization compiler [5][6], and has since expanded into a hosted inference platform with a serverless model API, dedicated GPU deployments and an early-access release of its compiler [0][3].
Risks & controversies
Public data on the company is inconsistent: third-party profiles variously report 30 employees and 11-20 employees, and one aggregator states the company has never raised funding despite the announced 2023 pre-seed round, while its revenue and valuation figures are labelled estimates derived from industry averages rather than disclosed results [5][7].
Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Timeline · 2
launches, deals, and filingsThe company opened early access to its compiler, which auto-compiles models for efficiency on target hardware.
Inceptron announced a pre-seed round reported as $2.2M (€2 million, USD 2.12m) led by 42CAP and Dreamcraft Ventures, to develop an optimization compiler for AI model deployment.
$2.1M source ↗
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
▸Research sources · 8
primary sources listed
- Inceptroninceptron.io · web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Inceptron do?
- Swedish AI inference platform offering serverless model APIs and a proprietary optimization compiler for production deployments.
- Who founded Inceptron?
- Inceptron was founded by Steffen Malkowsky in 2021.
- Who are Inceptron's investors?
- Inceptron's investors include 42CAP, Insiders Ventures, LU Innovation, PROfounders Capital, Dreamcraft.
