Your AI, in Spain and your prompts are never stored
Private LLM inference with an OpenAI-compatible API, on NVIDIA GPUs in our Cogent data center in Madrid. Your prompts are not stored or used for training, and everything is processed in Spain.
- OpenAI-compatible API
- 576 GB of VRAM
- Prompts not stored
- GDPR, data in Spain
Switching is one line
If you already use the OpenAI SDK, point base_url at our endpoint and pick the model. Chat completions and streaming, same as today.
- chat/completions
- stream: true
- Pick the model by name
1from openai import OpenAI23client = OpenAI(4 base_url="https://api.hostealo.com/v1",5 api_key="YOUR_API_KEY",6)78stream = client.chat.completions.create(9 model="qwen3-32b",10 messages=[{"role": "user", "content": "Summarize this contract in 3 points"}],11 stream=True,12)13for chunk in stream:14 print(chunk.choices[0].delta.conExample endpoint: we give you yours when the service is activated.
Two GPU servers, each with its own job
One server of each type in the Cogent data center in Madrid. We pick where your model runs based on its size, quantization and the kind of load.
4x NVIDIA RTX PRO 6000 Blackwell
384 GB of VRAM in a single server for large models: the model is split across the four GPUs and answers from one API.
- GPU
- 4x NVIDIA RTX PRO 6000 Blackwell
- VRAM per GPU
- 96 GB
- Total VRAM
- 384 GB
- Architecture
- Blackwell
- Location
- Cogent, Madrid
Built for large models such as GPT-OSS 120B or Llama 3.x 70B.
4x NVIDIA RTX 6000 Ada
192 GB of VRAM for many requests at once: mid-size models, embeddings and transcription, with requests grouped into batches and spread across the four GPUs.
- GPU
- 4x NVIDIA RTX 6000 Ada
- VRAM per GPU
- 48 GB
- Total VRAM
- 192 GB
- Architecture
- Ada Lovelace
- Location
- Cogent, Madrid
Built for 7B to 32B models, embeddings, Whisper and high-concurrency loads.
The models you already use
Common open families, served through the same API. Which size fits and on which server depends on the model and its quantization: we confirm it before activation.
Qwen 2.5 / Qwen3
DeepSeek R1 (distill)
Mistral / Mixtral
Gemma
GPT-OSS
Phi
Embeddings
Whisper
Your model or a custom solution
We load your own or fine-tuned weights and build complete solutions: RAG over your private documents, agents and internal assistants.
Depending on size and quantization. No guaranteed performance figures.
Prompt in, answer out, nothing kept
We process your request in Madrid and forget it. No content logs, no training on your data and nothing leaves Spain.
Prompts are not stored
Neither prompts nor answers are kept.
No training on your data
Your requests are never used to train models.
Everything in Spain
Inference runs on our servers at Cogent in Madrid.
GDPR
Data processed in Spain, inside the European Union.
Anti-DDoS in front
The service sits behind Hostealo Shield, our mitigation in Madrid.
Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category).
What companies use it for
Cases that fit a private AI in Spain.
Internal assistant over your documents
Customer support
Document analysis
Public sector and ENS
Developers
Transcription and search
Frequently asked questions
What people usually ask before requesting access.
It is a large language model (LLM) inference service on our GPU servers in Madrid, with an OpenAI-compatible API. Your applications send requests to our endpoint and the models run in Spain, on Hostealo hardware.
Yes. The chat completions endpoint is compatible, streaming included. With the official OpenAI SDK you only change base_url and the API key, and pick the model by name.
No. Prompts and answers are not stored and are never used to train models. The request is processed and discarded.
On our servers in the Cogent data center in Madrid. All processing happens in Spain, inside the European Union, in line with the GDPR.
Common open families such as Llama 3.x, Qwen 2.5 and Qwen3, the DeepSeek R1 distills, Mistral and Mixtral, Gemma, GPT-OSS 20B and 120B and Phi, plus embedding models (BGE, E5, Qwen) and Whisper for transcription. Which size fits depends on the model and its quantization; we confirm it before activation.
Yes. We load your own or fine-tuned weights as long as they fit in our GPU memory, and we also build custom solutions such as RAG over your private documents or agents.
Two servers in Madrid: one with 4x NVIDIA RTX PRO 6000 Blackwell 96 GB (384 GB of VRAM) for large models and one with 4x NVIDIA RTX 6000 Ada 48 GB (192 GB of VRAM) for mid-size models, embeddings and high-concurrency loads. 576 GB of VRAM in total.
We do not publish prices: we prepare a proposal based on the use case, model and volume. Reach us on WhatsApp, by email or open a request from your client area.
Access is given on request: there is no self sign-up and no free plan. Tell us what you want to build and we will explain how to get started.
Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category), and data is processed in Spain. If your project has specific requirements, tell us when you request access and we will review them with you.
Tell us what you want to build
There is no self sign-up and no public pricing: we set up access based on your use case, model and volume. Reach us on the channel you prefer.
If you can, tell us: company, use case, model you are interested in and approximate volume.

