Fundraising Fox

Inceptron

Founded 2021 · 20 employees on LinkedIn · 6 known investors

Inceptron provides a platform for deploying, optimizing, and scaling AI models in production, offering compiler-driven performance optimization and inference infrastructure for machine learning teams.

Also known as Inceptron AB

Founders & leadership

Inceptron was founded in 2021 by Steffen Malkowsky.

SMSteffen Malkowsky
Steffen MalkowskyinFounder

Investors · 6

Also in the syndicate · 1

Dreamcraft Ventureslead

Company profile

researched Aug 2026

Inceptron operates a platform for running AI models in production, combining serverless inference endpoints, dedicated GPU deployments and a proprietary optimization compiler. Users can call open-source or fine-tuned models through an OpenAI-compatible API at api.inceptron.io, import custom checkpoints, or start from a curated model library, then create versioned endpoints with API keys, access controls and rollout policies. The hosted catalogue includes large open-weight models such as GLM-5.2, Kimi-K2.7 Code, Kimi-K2.6, DeepSeek-V4 Flash and MiniMax-M2.5, listed with per-token pricing, quantization format, parameter size, context length and throughput figures, alongside Llama and Qwen family models.

The optimization layer is built around a compiler that fuses computation graphs, auto-tunes kernels and plans memory for a target hardware configuration, plus performance-aware compression through quantization and pruning. Company material describes automated passes including shift-based reparameterization, mixed-precision optimization, memory and cache tuning, and workload-specific search, applied across GPUs, CPUs and accelerators such as AWS Inferentia and FPGAs, for training, diffusion-model deployment and edge or cloud inference. Operational features include dynamic request batching, autoscaling, scheduled inference, integrated logging and tracing, and connections to cloud storage buckets, CI/CD and external observability stacks.

The platform is positioned for enterprise use with team controls, encryption and hardened isolation, ISO 27001 and GDPR compliance claims, data-residency controls including EU-only residency, SLA-backed uptime and zero-retention data handling.

Founding story

The founding team came from machine-learning research backgrounds, with experience described as spanning AI labs and infrastructure companies, and built an optimization compiler aimed at making model deployment across varied hardware and frameworks more efficient [5][7]. Named co-founders include Steffen Malkowsky, listed as CTO and founder, and a co-founder identified as Anton M. [5][7].

Business model

Inceptron sells access to its inference platform directly to engineering and machine-learning teams, with self-serve serverless usage alongside contracted dedicated deployments and enterprise add-ons such as custom regions, private networking and premium support quoted separately [0][4].

Serverless inference is billed per token, with separate input and output token rates rounded to the nearest 1,000 tokens per request batch (for example Llama-3.3-70B Instruct at $0.10 per million input and $0.30 per million output tokens); dedicated deployments are billed hourly per GPU pro-rated to the minute (for example an H100 80 GB instance at $3/hour). Commitment discounts are offered at 5% for one month, 10% for six months and 20% for twelve months, applying to both dedicated and serverless usage, with overage charged at on-demand rates and unused prepaid amounts rolling over within the commitment window [4]. Published catalogue prices range from $0.13 per million input tokens for DeepSeek-V4 Flash to $1.20 for GLM-5.2 [3].

Traction

Reported headcount is in the 11-20 range, with employees across engineering, design and sales functions [7]. The compiler is in an early-access phase and the hosted platform serves a catalogue of open-weight models with published pricing and throughput figures [0][3].

Latest developments

The compiler has been opened for early access with automated model compilation, and new models including DeepSeek V4 Flash have been added to the serving catalogue [0][3].

Full profile — market position, technology, go-to-market, geography, history, risks & controversies

Market position

Positions itself on price-performance for AI inference, offering hosted open-weight models at published per-token rates and pitching compiler-driven optimization as the source of lower latency and cost relative to unoptimized serving [0][3].

Company material emphasises a compiler that is applied to models without usage restrictions or added overhead, automated benchmarking and compile-time profiling, accuracy-preserving or use-case-tuned optimization delivered as a modular system, and engineering support alongside the hosted platform [0][7].

Technology

A proprietary inference-optimization compiler performs graph-level fusion, hardware-aware compilation, kernel auto-tuning and memory planning, combined with model compression via quantization and pruning. Auto-tuning uses ML agents together with Bayesian optimization to search for efficient algorithm implementations for a given model and target hardware, with tuning results aggregated in databases and as model weights to improve subsequent tuning. The serving layer adds dynamic batching, autoscaling, scheduled inference and unified observability correlating logs, metrics and traces across functions, containers and workloads, and exposes an OpenAI-compatible chat completions API [0][2]. Optimization targets include GPUs, CPUs and accelerators such as AWS Inferentia and FPGAs, with support for hybrid clouds and VPCs [7].

Go-to-market

Self-serve sign-up and API access on the website, a documentation site with quick-start code samples, a public model catalogue with a playground, an early-access program for the compiler, a Discord community for support and feature requests, and a sales contact form routing enquiries by use case (custom deployment, serverless inference, model/hardware optimization, fine-tuning) [0][2][3][4].

Machine-learning and product engineering teams deploying open-source or fine-tuned models in production, spanning prototyping and workload evaluation through to enterprise buyers requiring SLA-backed uptime, EU data residency and custom deployments [0][3][4].

Geography

Headquartered in Sweden, at Scheelevägen 15, Alpha 3 (Ideon Agora), Lund, Skåne County [5][7]. The platform offers region selection and EU-only data residency for compute and logs [3][4].

History

Founded in 2021 in Sweden [5][7]. The company announced a pre-seed round in November 2023 while developing an AI optimization compiler [5][6], and has since expanded into a hosted inference platform with a serverless model API, dedicated GPU deployments and an early-access release of its compiler [0][3].

Risks & controversies

Public data on the company is inconsistent: third-party profiles variously report 30 employees and 11-20 employees, and one aggregator states the company has never raised funding despite the announced 2023 pre-seed round, while its revenue and valuation figures are labelled estimates derived from industry averages rather than disclosed results [5][7].

Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Commitment discount - 12 monthsJan 202520%
Dedicated deployment price - H100 80GBJan 2025$3
HeadcountJan 202511-20
Serverless price - DeepSeek-V4 Flash 0731 input tokensJan 2025$0.13
Serverless price - DeepSeek-V4 Flash 0731 output tokensJan 2025$0.28
Serverless price - GLM-5.2 input tokensJan 2025$1.2
Serverless price - GLM-5.2 output tokensJan 2025$4.2

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Timeline · 2

launches, deals, and filings
Jan 2025
Inceptron compiler opened for early access

The company opened early access to its compiler, which auto-compiles models for efficiency on target hardware.

source ↗

Nov 2023
Inceptron announces pre-seed round for AI optimization compiler

Inceptron announced a pre-seed round reported as $2.2M (€2 million, USD 2.12m) led by 42CAP and Dreamcraft Ventures, to develop an optimization compiler for AI model deployment.

$2.1M source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

Research sources · 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Inceptron do?
Swedish AI inference platform offering serverless model APIs and a proprietary optimization compiler for production deployments.
Who founded Inceptron?
Inceptron was founded by Steffen Malkowsky in 2021.
Who are Inceptron's investors?
Inceptron's investors include 42CAP, Insiders Ventures, LU Innovation, PROfounders Capital, Dreamcraft.