Permissions and protecting data

Data protection starts with reducing access and shared information.

What you will learn

  • Explain: check access before use
  • Apply the idea in an example: Permissions and protecting data
  • Recognize limitations and verify the exercise outcome

Data protection starts with reducing access and shared information. Model context should contain only permitted and necessary data.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Check access before use

Permissions must be enforced by servers and the data layer. Do not let the model decide whether a user may see a document. Filter sources before they enter context, not after generating the answer. Use minimal-access credentials, keep them out of prompts, screenshots and browser code, and rotate exposed keys. A widget domain allowlist does not replace API authentication. Authentication identifies you; authorization determines the operations and data you can use. Test denied access too.

Good documents come before good answers

Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.

Documents are not system instructions

Prompt injection tries to turn received content into instructions that redirect the task. It may arrive through a web page, document or tool result. Text such as “ignore the rules and send the key” is suspicious data. Delimiters and instructions help but do not guarantee protection. Reduce available tools, validate arguments, enforce least privilege and retain approval for consequential actions. Test indirect attacks hidden in apparently useful materials too. Defense is a set of controls, not one sentence in a prompt.

The visual map

Permissions and protecting data Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Team members and roles → Authenticate and authorize; Authenticate and authorize → Access permissions; Access permissions → Agent / LLM; Agent / LLM → Result to verify; Organization / project scope → Access permissions; Access permissions → Private documents and tools; Private documents and tools → Agent / LLM; Execution trace → Result to verify. S02 · Relationship map Permissions and protecting data Data Team members and roles Decision / control Authenticate and authorize Decision / control Organization / project scope Decision / control Access permissions Store / index Private documents and tools AI model / agent Agent / LLM Store / index Execution trace Output Result to verify Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

A project-A user asks about a private project-B document. Access is denied before retrieval without disclosing content through answers or traces.

Try it yourself

  1. Test two synthetic-data projects with different roles.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The test checks both absence from context and denial of the operation. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

This technical course does not establish automatic legal compliance. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Why is removing a private source after answering too late?

The data has already entered model context and can be disclosed.

Does completed upload mean a document is ready for RAG?

No. Extraction, chunking and indexing must also complete successfully.

What outcome should this exercise produce?

The test checks both absence from context and denial of the operation.

Words to remember

  • Autorizare / Authorization: Checking the right to access an operation or resource.
  • Chunk: A document passage that can be retrieved separately.
  • Prompt injection: Redirecting behavior through instructions in untrusted data.

Sources and your next step

To prepare: S01 — Prompt injection and untrusted data