Services · LLM inference
A growing catalogue of AI models via API - and your data stays in Europe.
We run a growing catalogue of models in our infrastructure - from the flagship DeepSeek V4 Flash (a generative model) to Qwen3-Embedding-8B (an embedding model quantised to Q8). You connect once through an API and pay only for the tokens you use. No sending data outside the EU, no storing of prompts, no dependence on external vendors.
Flagship DeepSeek V4 Flash · Qwen3-Embedding-8B embeddings (Q8) · Context up to 1M tokens · Data in the EU
Models on offer
A model catalogue that keeps growing.
We serve different models for different tasks - we pick the one that fits your scenario best.
| Model | Type | Purpose | Context | Quantisation | Status |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | Generative LLM | Chat, agents, content generation and analysis | up to 1M tokens | FP8 | Available |
| Qwen3-Embedding-8B | Embedding model | Semantic search, RAG, classification | 32k tokens | Q8 | Available |
The catalogue keeps growing - more generative and embedding models are planned. Tell us what you need and we will pick the right model.
Why API
LLM inference, without infrastructure costs.
A premium-class model as a service - you connect through an API and pay for the tokens you use. We take care of the rest.
Your data stays in the EU
Prompts and responses never leave Europe. This matters for companies covered by GDPR that process personal, financial or medical data.
A flagship global-class model, focused on Polish
DeepSeek V4 Flash - quality comparable to the strongest models from major vendors and full proficiency in Polish. Plus embedding models for search and RAG.
You pay for tokens, not GPUs
No buying or maintaining infrastructure. You start from zero tokens and scale up as usage grows.
Zero prompt storage
We do not train on your data and we do not store conversations. Full control and transparency.
No vendor lock-in
Open models under MIT and Apache 2.0 licences. At any time you can move the workload to your own hardware - we will deploy it for you, without rewriting integrations.
Performance, not just quality
Context of up to a million tokens and fast responses thanks to advanced speculative decoding.
The model up close
What DeepSeek V4 Flash can do.
Matches the strongest.
Agent-task results comparable with the strongest premium-class models - e.g. Terminal Bench 2.1: 82.7, Cybergym: 76.7, Toolathlon: 70.3 - at a much lower cost.
Thinks as much as needed.
Three reasoning-depth modes (low / high / max) let you match the thinking effort to the task - simple queries are fast and cheap, while complex problems get full deliberation.
Remembers whole documents.
Context of up to a million tokens - you can provide entire books, legal codes, specifications or conversation archives without truncation.
Works in your company.
It connects beautifully with tools and completes tasks - an ideal foundation for automations and assistants that actually work.
Two types of models
Generative vs embedding - what each one does.
These are not two variants of the same model but two different categories that complement each other beautifully.
DeepSeek V4 Flash
It takes text and produces text. It answers questions, writes, edits, reasons and performs agentic tasks. It is the foundation of chat, assistants and content generation.
- Chat and assistants
- AI agents and automation
- Content generation and editing
- Document analysis and summaries
Qwen3-Embedding-8B
It takes text and turns it into a numeric vector describing its meaning. It does not generate text - it is used to find similar and relevant content in your data.
- Semantic search
- RAG - context retrieval
- Classification and clustering
- Similarity matching
Together they make RAG
Qwen3-Embedding-8B finds the most relevant fragments in your data, and DeepSeek V4 Flash generates a coherent answer in Polish based on them. It is the classic, proven RAG pattern.
Use cases
Where the model pays off fastest.
The scenarios we most often connect to the API - from assistants to process automation.
Assistants and chatbots
In Polish, running on your company data - customer service, helpdesk, internal support.
AI agents
The model completes tasks on its own: it works in a terminal, writes and tests code, and automates processes.
Analysis and summaries
With a context of up to 1M tokens you can process whole documents, reports and contracts - without splitting them into fragments.
Content generation and editing
Descriptions, offers, documents, correspondence - in Polish, in the tone of your company.
Integration with your systems
CRM, ERP, document workflow - through APIs and process automation.
Search and RAG
The embedding model finds relevant fragments in your data, and the LLM generates a coherent answer based on them - a ready-made RAG chain.
No obligations
Try LLM inference before you decide on the scale.
We will send you a test key - you will connect in minutes and see the quality on your own data. You only pay for actual token usage, and whenever you want, we will deploy the model in your infrastructure too.
How we work
From a test key to production.
Scenario selection
Together we identify where the model will bring the most value - chat, agent or document processing.
Integration
We connect your application or system to the API and configure the reasoning mode and limits.
Deployment
You go live in a production environment. You get API access and usage monitoring.
Growth and care
We fine-tune parameters, track costs, train your team and extend scenarios.
Specification
The models at a glance.
DeepSeek V4 Flash
Generative LLM
- Model name
- DeepSeek V4 Flash
- Type
- Generative LLM
- Context
- up to 1 million tokens
- Reasoning modes
- low / high / max
- Licence
- MIT (open weights)
- Quantisation
- FP8
- API
- OpenAI-compatible, REST interface
- Access
- via NetCreate infrastructure - data never leaves the EU
- Languages
- Polish, English and others
Agentic benchmarks · per the model card
82.7
Terminal Bench 2.1
76.7
Cybergym
70.3
Toolathlon-Verified
54.4
DeepSWE
54.2
NL2Repo
Qwen3-Embedding-8B
Embedding model · Q8
- Model name
- Qwen3-Embedding-8B
- Type
- Embedding model (text vectorisation)
- Parameters
- 8B
- Context
- 32k tokens
- Vector dimension
- up to 4096 (configurable 32-4096, MRL)
- Languages
- 100+ (including Polish and programming languages)
- Licence
- Apache 2.0 (open weights)
- Quantisation
- Q8
- API
- /embed endpoint, REST interface
- Access
- via NetCreate infrastructure - data never leaves the EU
MTEB benchmarks · per the model card
70.58
MTEB Multilingual
75.22
MTEB English v2
70.88
Retrieval (ML)
81.08
STS (ML)
73.84
C-MTEB
FAQ
The questions we hear most often.
Contact
Let us talk about your project.
Tell us about your idea - we will respond, advise and quote the project. No obligations.
ul. Dworcowa 11B, 05-820 Piastów
NIP 534-219-20-63 · REGON 140520107