How to NOT Have Your Data Trained on by AI Across Leading LLMs

Large Language Models (LLMs) have become indispensable tools for businesses, but a key concern persists: How do you ensure your private or proprietary data isn’t used to train these models? This guide covers the most notable LLMs, their policies, and actionable steps to safeguard your data. Links to turn off data training for each tool are provided for easy access.

GPT-4: OpenAI

Notable Features: 175 billion parameters; widely used in enterprise applications.

How OpenAI Handles Data:

Default Behavior: ChatGPT logs conversations for improving its systems unless explicitly disabled.

Data Training: Conversations are not used for training the model when API services are used.

How to Protect Your Data:

  1. Use the API Instead of Public Chat: Data processed via OpenAI’s API (e.g., Azure OpenAI Service) is excluded from training datasets.
    Review OpenAI's API Privacy Policy
  2. Turn Off Chat History: For the free or premium ChatGPT app, disable conversation history in settings.
    Turn Off Chat History
  3. Review Privacy Policies: Opt for Enterprise or API versions, which have stricter data handling practices.

Use OpenAI's Enterprise API or ensure history is turned off for sensitive conversations.

Claude 3.5: Anthropic

Notable Features: Focuses on constitutional AI to produce safer and more accurate responses.

How Anthropic Handles Data:

Anthropic collects user inputs to improve safety but claims inputs are not used for training unless you explicitly opt-in.

API users’ data is excluded from model training.

How to Protect Your Data:

  1. Opt Out: Avoid sending sensitive data in casual conversations. For API users, ensure you verify that your usage agreements include data exclusion from training.
    Anthropic Privacy Policy
  2. Contact Support for Clarification: If needed, request written confirmation of your data protection policies.
    Contact Anthropic Support

Check Anthropic’s Data Privacy Agreement for enterprise customers to confirm exclusions.

Gemini: Google DeepMind

Notable Features: Multimodal capabilities integrated into Google products.

How Google Handles Data:

Conversations through Gemini (formerly Bard) may be analyzed for quality and training unless users opt out.

Data sent via Google APIs, particularly for Google Cloud customers, is not used for training.

How to Protect Your Data:

  1. Leverage Google Cloud Enterprise API: Use enterprise-grade APIs for projects that require sensitive data handling.
    Google Cloud Privacy Policy
  2. Disable History: Turn off activity tracking on Gemini products where possible.
    Manage Activity Settings
  3. Verify Privacy Settings: Use tools like Google Admin to manage data retention policies.

Use Google Cloud APIs for commercial LLM projects to ensure strict privacy controls.

Llama 3.1: Meta

Notable Features: Open-source; model sizes from 8B to 405B parameters.

How Meta Handles Data:

Since Llama 3.1 is open-source, Meta does not collect your data unless you interact with Meta’s hosted services.

Responsibility lies with the user to host the model securely.

How to Protect Your Data:

  1. Self-Host: Deploy Llama models on-premise or in secure cloud environments to avoid data exposure.
    Access Llama Models
  2. Review Meta's Privacy Policy: For Meta-hosted services, review their policies.
    Meta Privacy Policy

Host Llama locally or in a private cloud environment to fully control access.

Mistral 7B: Mistral AI

Notable Features: Compact model size with high efficiency.

How Mistral Handles Data:

Mistral’s models are open-source, so they don’t collect or process user data.

Security and data privacy depend on your hosting strategy.

How to Protect Your Data:

  1. Deploy Privately: Use private servers or self-host to maintain full control over data flows.
    Access Mistral Models
  2. Restrict API Access: If third-party APIs are used, ensure they don’t collect input/output data.
    Mistral GitHub Repository

Secure your model deployment with role-based access controls and encryption.

Falcon 180B: Technology Innovation Institute

Notable Features: Open-source; excellent benchmark performance.

How TII Handles Data:

Falcon is open-source, meaning no data is collected by the developer.

As with Llama, user security is dependent on how the model is deployed.

How to Protect Your Data:

  1. Run Falcon Locally: Use private infrastructure for deployment.
    Access Falcon Models
  2. Audit Logs: Ensure activity logs are monitored to detect unauthorized access.
    Falcon License Agreement

Implement secure hosting practices like Kubernetes isolation or VPC deployment.

Cohere

Notable Features: Customizable models; enterprise-grade multilingual support.

How Cohere Handles Data:

Data used via the API is not stored or used for training.

Cohere’s Enterprise API offers explicit guarantees for privacy and data protection.

How to Protect Your Data:

  1. Use the Enterprise API: Opt for enterprise-grade solutions to ensure compliance with data protection regulations.
    Cohere Privacy Policy
  2. Request Custom Policies: Work with Cohere to define custom privacy agreements if needed.
    Contact Cohere

Confirm your organization’s usage aligns with Cohere’s enterprise data policies.

Grok-1: xAI

Notable Features: Integrates with X (formerly Twitter); open-source with a massive parameter size.

How xAI Handles Data:

Grok-1 is open-source. Data protection depends on user-hosted environments.

Integration with X introduces potential risks of data logging through APIs.

How to Protect Your Data:

  1. Host the Model Yourself: Avoid relying on X-integrated services.
    Access Grok on GitHub
  2. Monitor API Usage: Scrutinize any API connections that could inadvertently send data to external systems.
    xAI Privacy Policy

Isolate Grok-1 deployments on secure servers with limited network exposure.

DBRX: Databricks

Notable Features: Combines open-source accessibility with enterprise capabilities.

How Databricks Handles Data:

Databricks ensures data processed on their platform is not used for training without user consent.

The platform adheres to enterprise-grade compliance standards.

How to Protect Your Data:

  1. Leverage Databricks Enterprise: Use enterprise-grade solutions to ensure compliance.
    Databricks Data Privacy
  2. Review Agreements: Confirm policies exclude your datasets from model retraining.
    Contact Databricks Support

Conduct a Data Privacy Audit on Databricks agreements for peace of mind.

Microsoft Copilot

Notable Features: AI-powered tools integrated into Microsoft 365, enhancing productivity with real-time AI assistance.

How Microsoft Handles Data:

Microsoft Copilot processes data in accordance with enterprise-grade privacy standards.

Microsoft ensures data processed through Copilot is not used for training without explicit consent.

Microsoft provides data encryption and compliance with GDPR and other regulations.

How to Protect Your Data:

  1. Use Microsoft Enterprise Services: Deploy Copilot through Microsoft 365 Enterprise, which ensures robust data privacy.
    Azure OpenAI Privacy Policy
  2. Review Privacy Settings: Use Microsoft Admin tools to manage permissions and audit usage.
    Azure Trust Center

Configure Microsoft 365 settings to align with your organization’s compliance requirements.

Microsoft Orca

Notable Features: Optimized for performance with fewer parameters.

How Microsoft Handles Data:

Data processed via Microsoft Azure services (including OpenAI GPT) is not used for training.

Microsoft provides clear guarantees for enterprise data security.

How to Protect Your Data:

  1. Choose Azure OpenAI Services: Azure integrates stringent compliance and privacy controls.
    Azure OpenAI Privacy Policy
  2. Review Compliance Reports: Use Microsoft’s security and compliance dashboards to confirm data usage.
    Azure Trust Center

Opt for Microsoft Azure-hosted LLMs with data encryption and user access control.

Key Takeaways for Businesses

  1. Self-Host Open-Source Models: Models like Llama, Mistral, and Falcon offer the greatest control if hosted securely.
  2. Use Enterprise APIs: Services like OpenAI, Cohere, and Microsoft Azure provide contractual guarantees for data privacy.
  3. Review Agreements Regularly: Ensure your provider explicitly excludes your data from model training.
  4. Implement Encryption: For any hosted model, encrypt data at rest and in transit to prevent leaks.

By following these strategies and leveraging the provided resources, businesses can confidently integrate AI while ensuring their proprietary data remains private and secure. Ready to deploy an LLM securely? Contact Facet today to explore customized AI solutions for your business.