Skip to content
New · Madrid, Spain

Your AI, in Spain and your prompts are never stored

Private LLM inference with an OpenAI-compatible API, on NVIDIA GPUs in our Cogent data center in Madrid. Your prompts are not stored or used for training, and everything is processed in Spain.

  • OpenAI-compatible API
  • 576 GB of VRAM
  • Prompts not stored
  • GDPR, data in Spain
OpenAI API

Switching is one line

If you already use the OpenAI SDK, point base_url at our endpoint and pick the model. Chat completions and streaming, same as today.

  • chat/completions
  • stream: true
  • Pick the model by name
1from openai import OpenAI
2
3client = OpenAI(
4 base_url="https://api.hostealo.com/v1",
5 api_key="YOUR_API_KEY",
6)
7
8stream = client.chat.completions.create(
9 model="qwen3-32b",
10 messages=[{"role": "user", "content": "Summarize this contract in 3 points"}],
11 stream=True,
12)
13for chunk in stream:
14 print(chunk.choices[0].delta.con
Output (streaming)Response complete
1. 24-month term with automatic renewal. 2. 20% penalty for early termination. 3. Data is processed in the EU under the GDPR.

Example endpoint: we give you yours when the service is activated.

Hardware in Madrid

Two GPU servers, each with its own job

One server of each type in the Cogent data center in Madrid. We pick where your model runs based on its size, quantization and the kind of load.

576 GB VRAM8 GPU NVIDIACogent Madrid
Server 1384 GB

4x NVIDIA RTX PRO 6000 Blackwell

384 GB of VRAM in a single server for large models: the model is split across the four GPUs and answers from one API.

GPU
4x NVIDIA RTX PRO 6000 Blackwell
VRAM per GPU
96 GB
Total VRAM
384 GB
Architecture
Blackwell
Location
Cogent, Madrid

Built for large models such as GPT-OSS 120B or Llama 3.x 70B.

Server 2192 GB

4x NVIDIA RTX 6000 Ada

192 GB of VRAM for many requests at once: mid-size models, embeddings and transcription, with requests grouped into batches and spread across the four GPUs.

GPU
4x NVIDIA RTX 6000 Ada
VRAM per GPU
48 GB
Total VRAM
192 GB
Architecture
Ada Lovelace
Location
Cogent, Madrid

Built for 7B to 32B models, embeddings, Whisper and high-concurrency loads.

The models you already use

Common open families, served through the same API. Which size fits and on which server depends on the model and its quantization: we confirm it before activation.

Llama 3.x

8B, 70B
chatcodemultilingual

Qwen 2.5 / Qwen3

7B to 72B
chatcodereasoning

DeepSeek R1 (distill)

7B to 70B
reasoning

Mistral / Mixtral

7B, 24B, 8x7B
chatmultilingual

Gemma

2B to 27B
chatsummaries

GPT-OSS

20B, 120B
chatreasoningagents

Phi

3.8B, 14B
chatlightweight

Embeddings

BGE, E5, Qwen
searchRAG

Whisper

speech to text
transcription

Your model or a custom solution

We load your own or fine-tuned weights and build complete solutions: RAG over your private documents, agents and internal assistants.

Tell us about it

Depending on size and quantization. No guaranteed performance figures.

RGPD

Prompt in, answer out, nothing kept

We process your request in Madrid and forget it. No content logs, no training on your data and nothing leaves Spain.

  • Prompts are not stored

    Neither prompts nor answers are kept.

  • No training on your data

    Your requests are never used to train models.

  • Everything in Spain

    Inference runs on our servers at Cogent in Madrid.

  • GDPR

    Data processed in Spain, inside the European Union.

  • Anti-DDoS in front

    The service sits behind Hostealo Shield, our mitigation in Madrid.

Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category).

What companies use it for

Cases that fit a private AI in Spain.

Internal assistant over your documents

RAG over contracts, manuals or internal tickets without them leaving Spain.

Customer support

Answers and drafts for your support team, grounded in your own knowledge base.

Document analysis

Summaries, data extraction and classification of long texts.

Public sector and ENS

A Spanish provider with ISO 27001 and ENS medium category for projects with sensitive data.

Developers

Replace the OpenAI API by changing one line and keep using the same SDK.

Transcription and search

Whisper to turn audio into text and embeddings to search your information.

Frequently asked questions

What people usually ask before requesting access.

It is a large language model (LLM) inference service on our GPU servers in Madrid, with an OpenAI-compatible API. Your applications send requests to our endpoint and the models run in Spain, on Hostealo hardware.

Yes. The chat completions endpoint is compatible, streaming included. With the official OpenAI SDK you only change base_url and the API key, and pick the model by name.

No. Prompts and answers are not stored and are never used to train models. The request is processed and discarded.

On our servers in the Cogent data center in Madrid. All processing happens in Spain, inside the European Union, in line with the GDPR.

Common open families such as Llama 3.x, Qwen 2.5 and Qwen3, the DeepSeek R1 distills, Mistral and Mixtral, Gemma, GPT-OSS 20B and 120B and Phi, plus embedding models (BGE, E5, Qwen) and Whisper for transcription. Which size fits depends on the model and its quantization; we confirm it before activation.

Yes. We load your own or fine-tuned weights as long as they fit in our GPU memory, and we also build custom solutions such as RAG over your private documents or agents.

Two servers in Madrid: one with 4x NVIDIA RTX PRO 6000 Blackwell 96 GB (384 GB of VRAM) for large models and one with 4x NVIDIA RTX 6000 Ada 48 GB (192 GB of VRAM) for mid-size models, embeddings and high-concurrency loads. 576 GB of VRAM in total.

We do not publish prices: we prepare a proposal based on the use case, model and volume. Reach us on WhatsApp, by email or open a request from your client area.

Access is given on request: there is no self sign-up and no free plan. Tell us what you want to build and we will explain how to get started.

Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category), and data is processed in Spain. If your project has specific requirements, tell us when you request access and we will review them with you.

Access on request

Tell us what you want to build

There is no self sign-up and no public pricing: we set up access based on your use case, model and volume. Reach us on the channel you prefer.

If you can, tell us: company, use case, model you are interested in and approximate volume.