10 Security Loopholes Every Local AI User Should Know

When it comes to local AI, you might assume that running models on local hardware instead of paying for cloud services is a more secure and private option. This is only partially true. When running AI locally, you retain greater control over your workflows and data. You decide whether and under what conditions to share your data with third-party services. You also don’t have to worry about data leaks and cyberattacks targeting large tech corporations. But on the other hand, you’re giving up the multi-million-dollar security infrastructure built by established players like OpenAI or Anthropic. Your security is now entirely your responsibility, for better or for worse. With this in mind, here are ten helpful tips to help you improve security when running AI models on your PC or private VPS (virtual private server).
Is it safer to run AI models on local hardware?
When running AI models on your personal hardware using platforms like Jan, Ollama, or LM Studio, your messages, documents, and chat history never leave your device or are sent to third-party cloud servers. If you’re concerned about an AI company selling your data without your consent, or if you don’t want your credentials compromised in the next data breach, then a local approach is a smart choice.
However, this doesn’t guarantee complete protection. You’re still loading AI model files from a public database. You may also need to allow the model access to certain APIs so it can interact with your software or data over a public network. Finally, if you’re using a public Wi-Fi network, anyone connected to the network could hack your operating system by attacking the AI. In other words, things aren’t as simple as they sometimes seem.
In January of this year alone, SentinelOne and Censys discovered 175,000 publicly accessible Ollama hosts that could be exploited by any attacker with internet access to execute code and connect to third-party services using the user’s credentials and hardware. If you want to run AI locally, you need to be very careful about where you source your models and what they have access to. Here are some tips to help you do it right.
Host the model server on your local computer.
To run AI models locally, you’ll need to use an inference engine (also called an executor), such as Ollama or LM Studio, which allows you to load and execute the model on your hardware. By default, AI executors are configured to run models on localhost (127.0.0.1 or 0:0:0:0:0:0:0:1), meaning other devices on your network or the public internet can’t access them. However, running your AI model on 0.0.0.0 exposes it to all devices on your network. Anyone using your shared Wi-Fi network could run your local AI system and then use it to modify your hardware or steal sensitive data.
Sometimes setup guides suggest doing this anyway so you can access your local AI model from other devices on your network, such as a smartphone or laptop. This also happens when people try to run AI models on VPS servers or network-attached storage (NAS). However, this will compromise your data and workflows, so if you’ve changed the default server configuration for running your model, be sure to revert it back:
-
In Ollam this can be done by changing the OLLAMA_HOST variable back to 127.0.0.1.
-
In LM Studio, disable the “Run on local network” option.
-
If you’re using Jan, click the gear icon in the Hub interface to go to the settings page. Then select “Local API Server” and generate an API key using an online generator, such as RandomKeygen .
Instead of port forwarding, use a private VPN tunnel.
You shouldn’t expose your AI model to the public internet via your IP address. But what if you still need to share the model with other devices remotely? Port forwarding is typically enabled on routers to allow access to resources and data from a remote location. However, you should never use this approach to configure remote access to AI models or workers on your local computer.
If your AI model is already running on the IP address 0.0.0.0 and you additionally enable port forwarding on your router, any internet user will be able to hack your local AI system if they can guess your IP address. Cybercriminals often use botnets that regularly scan the internet for open ports on residential IP addresses, so you risk becoming a target if you do this. The best way is to set up an encrypted tunnel using a VPN or Cloudflare ZTNA . Mesh VPNs like Tailscale are popular for this, as is Cloudflare’s new Tunnel feature, Zero Trust.
Update your AI runner as soon as patches are released.
In May 2026, Cyera discovered a new vulnerability in Ollama that allowed attackers to steal fragments of your data and credentials using unauthenticated API calls. This vulnerability, dubbed ” Bleeding Llama ,” had a CVSS score of 9.3 out of 10. At the time, it compromised approximately 300,000 publicly accessible Ollama servers until it was patched in version 0.17.1.
AI-powered app launchers like LM Studio, Ollama, Jan, and GPT4All are still in the experimental phase and frequently discover new vulnerabilities that are patched in subsequent releases. If your app is even a few versions out of date, your server may be vulnerable to a serious attack that hackers can exploit. Always download the latest version as soon as possible from the app’s official website or GitHub repository.
Choose Safetensors or GGUF files instead of pickle.
AI models based on older deep learning models, such as PyTorch, can often be downloaded as pickle files with extensions like .bin, .pt, or .pkl. However, due to the nature of the Python pickle file format, these model files can be modified to execute malicious code when you attempt to load them using your launcher.
Back in 2025, ReversingLabs discovered two model files on the Hugging Face platform that passed the platform’s automatic security check despite the presence of a hidden feature for unauthorized remote access. Hugging Face’s documentation now considers pickle files a serious security threat.
To prevent data leaks or unauthorized access, only download LLM models in newer file formats, such as .safetensor or .gguf. These formats store your data numerically, making it impossible to execute malicious code when loading the model. If a specific model is only available in .pt or .pkl format, it’s best to skip it. There are many newer versions of LLM models that use more secure file formats.
Download models from publishers whose work you can check out.
AI platforms like Hugging Face or ModelScope allow anyone with internet access to upload AI models to their website. While they have some platform-level security protocols, in cases like the one discovered by ReversingLabs in 2025, newer or more complex vulnerabilities could easily bypass these protocols and checks.
To enhance security, upload model files only from official accounts managed by major model developers. For example, Google, Mistral, Meta, and Qwen (Alibaba) have separate corporate accounts with a verified badge on Hugging Face. These verified badges indicate that the company’s account is genuinely owned and managed by that company, as the user must use an official corporate email address to log in and upload model files. More information on how verified badges for enterprises work can be found in the “Advanced Security” section of the Hugging Face documentation.
Download AI-powered apps only from official websites.
Hackers like to use popular GenAI tools as bait to trick people into installing malicious software. They often create fake websites or upload apps to popular app stores, where they can impersonate official platforms. Early attempts focused on ChatGPT clones on websites resembling the real OpenAI. They tricked people into downloading a corrupted .exe or .dmg file, which then distributed dangerous malware such as Redline , Lumma, or the Odyssey infostealer for Mac . Similar attempts have also been used to attack Android users via malicious apps uploaded to the Play Store.
However, since then, attacks have become more sophisticated and can even target obscure local AI platforms and Python packages. Attackers have gone so far as to compromise official GitHub repositories and uploads to the Python Package Index (PyPI). TrendAI reported one particularly alarming case in which malicious code was injected directly into the official PyPI package of LiteLLM, an open-source AI gateway that allows calling hundreds of LLMs from a single API. Positive Security also discovered malicious Python packages uploaded to PyPI disguised as Deepseek packages.
Be sure to check where you’re getting your AI tools from. It’s best to rely on direct official sources, verified GitHub repositories maintained by trusted AI vendors, and Python packages directly referenced in the vendor’s official documentation.
Double-check the packages that your model recommends installing.
I’ve already discussed how Python packages are corrupted to install malware immediately after being executed on a system. But it’s not just LLM files and AI tools that need to be wary. When you ask AI agents to write code or perform tasks, they also install and run any packages or dependencies required to complete that work. And because AI models are prone to hallucinations, agents often simply invent package names that don’t exist.A recent study analyzing 16 models based on 576,000 code samples found that open-weight LLM models do this 21.7% of the time, while state-of-the-art AI models perform this task at a lower rate of 5.2%.
Hackers know this, hence “slopsquatting”—a new attack in which attackers register fake software packages under often fictitious names in various LLM systems. These packages can launch malicious code, trigger injection attacks, or data thefts once your AI agent executes them on your local computer.
The best way to avoid such attacks is to limit what your AI agent can install and run without your approval. You can either manually approve each software package before the model installs or runs it, or whitelist specific trusted repositories that are unlikely to contain malware. In any case, be sure to review your model’s logs to see which pip install and npm install commands it executes to avoid unauthorized installations.
Limit what your AI agents can touch.
Even when running on local hardware, AI agents can access MCP servers, download and run files, search the internet, or connect to third-party services via APIs. Furthermore, they can read and write files to local hardware and even modify key operating system settings. All these features should be enabled only with extreme caution, based on your security profile. Carefully manage the level of access an AI agent or model runner has to your system, especially for new open-source models, which are more likely to generate false positives or have exploitable vulnerabilities.
There are several ways to control access granted to an AI agent. The first is to run AI workflows within a Docker container , which prevents them from directly modifying system files. You can also restrict access rights by modifying the default configuration of your agent platform, such as OpenClaw or Hermes. OpenClaw allows you to choose between three default permission profiles: request, deny, and allow list, which can be further customized for specific workflows and services. Hermes also allows you to configure a similar list of allowed users (white list) or restrict tool usage for each cron job.
Enable local only mode.
AI model launch systems like Ollama and LM Studio support both local and cloud-based models. However, you can configure them to restrict network access using a single feature, even if you haven’t done so at the orchestration level with Hermes or OpenClaw. To do this, you need to bind the service to your local IP address (127.0.0.1) to prevent other devices from accessing it through your network or the public internet.
Encrypt the disk where your chat history is stored.
When you store your AI workflows locally, your entire chat history, as well as any credentials, secrets, or API tokens you may have shared with your model, are stored in cleartext on your local drives. If someone gains physical access to your device, they can retrieve all your data. Applications like FileVault , BitKocker , or LUKS can encrypt your hard drive so that your chat history cannot be read in cleartext without the encryption key to decrypt it. Use these to avoid the risk of data leakage if your device is stolen or lost.
Several local AI platforms to get you started.
If you’re new to local AI, here are a few platforms to experiment with. They provide the best access for new users unfamiliar with the technical aspects of AI development.
-
Ollama: An open-source model runner for macOS, Windows, and Linux. It features a large model library and a simple desktop application that can be connected to most other local AI tools.
-
LM Studio: An improved desktop application that allows you to download models from the Hugging Face website via a graphical interface. It’s free for both work and personal use since July 2025.
-
Jan: An open-source alternative to ChatGPT, licensed under the Apache 2.0 license, that runs completely offline on Windows, macOS, and Linux.
-
AnythingLLM Desktop: A free, MIT-licensed app for working with your own documents locally. It’s a great choice if you want to share PDFs and model notes without downloading them.
-
Open WebUI: a browser-based offline chat interface that can connect to Ollama, making the user interface more accessible. When combined with a mesh VPN, your entire family can securely use a single AI server.