# Your AI, in Spain and your prompts are never stored

Private LLM inference with an OpenAI-compatible API, on NVIDIA GPUs in our Cogent data center in Madrid. Your prompts are not stored or used for training, and everything is processed in Spain.

## Key facts

- New · Madrid, Spain
- OpenAI-compatible API
- 576 GB of VRAM
- Prompts not stored
- GDPR, data in Spain

## Switching is one line

If you already use the OpenAI SDK, point base_url at our endpoint and pick the model. Chat completions and streaming, same as today.

Example endpoint: we give you yours when the service is activated.

## Two GPU servers, each with its own job

One server of each type in the Cogent data center in Madrid. We pick where your model runs based on its size, quantization and the kind of load.

### Server 1: 4x NVIDIA RTX PRO 6000 Blackwell

384 GB of VRAM in a single server for large models: the model is split across the four GPUs and answers from one API.

| Spec |   |
| --- | --- |
| GPU | 4x NVIDIA RTX PRO 6000 Blackwell |
| VRAM per GPU | 96 GB |
| Total VRAM | 384 GB |
| Architecture | Blackwell |
| Location | Cogent, Madrid |

Built for large models such as GPT-OSS 120B or Llama 3.x 70B.

### Server 2: 4x NVIDIA RTX 6000 Ada

192 GB of VRAM for many requests at once: mid-size models, embeddings and transcription, with requests grouped into batches and spread across the four GPUs.

| Spec |   |
| --- | --- |
| GPU | 4x NVIDIA RTX 6000 Ada |
| VRAM per GPU | 48 GB |
| Total VRAM | 192 GB |
| Architecture | Ada Lovelace |
| Location | Cogent, Madrid |

Built for 7B to 32B models, embeddings, Whisper and high-concurrency loads.

## The models you already use

Common open families, served through the same API. Which size fits and on which server depends on the model and its quantization: we confirm it before activation.

| Family | Sizes | Use |
| --- | --- | --- |
| Llama 3.x | 8B, 70B | chat, code, multilingual |
| Qwen 2.5 / Qwen3 | 7B to 72B | chat, code, reasoning |
| DeepSeek R1 (distill) | 7B to 70B | reasoning |
| Mistral / Mixtral | 7B, 24B, 8x7B | chat, multilingual |
| Gemma | 2B to 27B | chat, summaries |
| GPT-OSS | 20B, 120B | chat, reasoning, agents |
| Phi | 3.8B, 14B | chat, lightweight |
| Embeddings | BGE, E5, Qwen | search, RAG |
| Whisper | speech to text | transcription |

Depending on size and quantization. No guaranteed performance figures.

### Your model or a custom solution

We load your own or fine-tuned weights and build complete solutions: RAG over your private documents, agents and internal assistants.

## Prompt in, answer out, nothing kept

We process your request in Madrid and forget it. No content logs, no training on your data and nothing leaves Spain.

- **Prompts are not stored:** Neither prompts nor answers are kept.
- **No training on your data:** Your requests are never used to train models.
- **Everything in Spain:** Inference runs on our servers at Cogent in Madrid.
- **GDPR:** Data processed in Spain, inside the European Union.
- **Anti-DDoS in front:** The service sits behind Hostealo Shield, our mitigation in Madrid.

Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category).

- [ISO 27001](https://hostealo.com/iso-27001-certificate.pdf)
- [ENS medium category](https://hostealo.com/ens-certificado-conformidad-medio.pdf)

## What companies use it for

Cases that fit a private AI in Spain.

- **Internal assistant over your documents:** RAG over contracts, manuals or internal tickets without them leaving Spain.
- **Customer support:** Answers and drafts for your support team, grounded in your own knowledge base.
- **Document analysis:** Summaries, data extraction and classification of long texts.
- **Public sector and ENS:** A Spanish provider with ISO 27001 and ENS medium category for projects with sensitive data.
- **Developers:** Replace the OpenAI API by changing one line and keep using the same SDK.
- **Transcription and search:** Whisper to turn audio into text and embeddings to search your information.

## Frequently asked questions

What people usually ask before requesting access.

### What is Hostealo private AI?

It is a large language model (LLM) inference service on our GPU servers in Madrid, with an OpenAI-compatible API. Your applications send requests to our endpoint and the models run in Spain, on Hostealo hardware.

### Is it compatible with the OpenAI API?

Yes. The chat completions endpoint is compatible, streaming included. With the official OpenAI SDK you only change base_url and the API key, and pick the model by name.

### Do you store my prompts or the answers?

No. Prompts and answers are not stored and are never used to train models. The request is processed and discarded.

### Where is the data processed?

On our servers in the Cogent data center in Madrid. All processing happens in Spain, inside the European Union, in line with the GDPR.

### Which models can I use?

Common open families such as Llama 3.x, Qwen 2.5 and Qwen3, the DeepSeek R1 distills, Mistral and Mixtral, Gemma, GPT-OSS 20B and 120B and Phi, plus embedding models (BGE, E5, Qwen) and Whisper for transcription. Which size fits depends on the model and its quantization; we confirm it before activation.

### Can you load my own or a fine-tuned model?

Yes. We load your own or fine-tuned weights as long as they fit in our GPU memory, and we also build custom solutions such as RAG over your private documents or agents.

### What hardware is behind it?

Two servers in Madrid: one with 4x NVIDIA RTX PRO 6000 Blackwell 96 GB (384 GB of VRAM) for large models and one with 4x NVIDIA RTX 6000 Ada 48 GB (192 GB of VRAM) for mid-size models, embeddings and high-concurrency loads. 576 GB of VRAM in total.

### How much does it cost?

We do not publish prices: we prepare a proposal based on the use case, model and volume. Reach us on WhatsApp, by email or open a request from your client area.

### Can I sign up and try it now?

Access is given on request: there is no self sign-up and no free plan. Tell us what you want to build and we will explain how to get started.

### Is it suitable for public sector projects?

Hostealo holds ISO 27001 certification and conformity with the Spanish ENS (medium category), and data is processed in Spain. If your project has specific requirements, tell us when you request access and we will review them with you.

## Tell us what you want to build

There is no self sign-up and no public pricing: we set up access based on your use case, model and volume. Reach us on the channel you prefer.

- **WhatsApp** (Quick answer from sales): [wa.me/34900433332](https://wa.me/34900433332?text=Hi%2C%20I%20am%20interested%20in%20Hostealo%20private%20AI%20(OpenAI-compatible%20API%20in%20Madrid).%20Can%20we%20talk%3F)
- **Email**: [contacto@hostealo.es](mailto:contacto@hostealo.es?subject=IA%20privada&body=Hi%2C%0A%0AI%20am%20interested%20in%20Hostealo%20private%20AI.%0A%0ACompany%3A%0AUse%20case%3A%0AModel%20we%20are%20interested%20in%3A%0AApproximate%20volume%20(requests%20or%20tokens%20per%20month)%3A%0A%0AThanks.)
- **Open a request** (From your client area): [my.hostealo.com/support/new](https://my.hostealo.com/support/new?subject=Access%20request%3A%20private%20AI&department=sales&body=Hi%2C%20I%20would%20like%20to%20request%20access%20to%20the%20private%20AI.%0A%0ACompany%3A%0AUse%20case%3A%0AModel%20we%20are%20interested%20in%3A%0AApproximate%20volume%20(requests%20or%20tokens%20per%20month)%3A)

If you can, tell us: company, use case, model you are interested in and approximate volume.

---

Canonical HTML version: https://hostealo.com/private-ai
