Skip to content
What we store and share

What we store and share

Teramot builds and operates the data warehouse of each project. To do so, it keeps a copy of the tables you choose to bring in, on Teramot’s own infrastructure in AWS. This page brings together in one place what information is stored, what information leaves that infrastructure, who receives it, and for how long.

Summary

    flowchart LR
    S[(Your systems)] -->|Read-only| X
    subgraph AWS["Teramot on AWS us-east-1"]
        X[Extractors] --> L[(Data lake<br/>AES-256 encrypted)]
        L --> Q[SQL engine]
        C[(Configuration database<br/>AES-256 encrypted)]
    end
    Q -->|Schemas, statistics,<br/>and samples| M[Model providers<br/>Anthropic, OpenAI]
    Q -->|Agent traces| T[AI observability<br/>LangSmith, Langfuse]
    X -->|Execution metadata| O[Orchestration<br/>Temporal Cloud]
  
  • Complete tables stay within Teramot’s infrastructure in AWS.
  • Model providers receive only schemas, statistics, and bounded samples, never complete tables.
  • Nothing is used to train models.

What we store

InformationWhereRetention
Raw data, processed data, and results tablesData lake on Amazon S3, Apache Iceberg formatFor as long as the source or project exists
Previous versions of each table (snapshots)Data lake on Amazon S33 days
Files uploaded by usersAmazon S3For as long as the source exists
Table schemasAWS Glue Data Catalog, one database per projectFor as long as the table exists
Configuration: workspaces, projects, members, sources, cleaning and results queries, dashboardsAmazon Aurora PostgreSQLFor as long as the object exists; backups for 7 days
Credentials for source systemsAmazon Aurora PostgreSQLFor as long as the source exists
Technical logsAmazon CloudWatch365 days

Encryption at rest

  • Data lake files, uploaded files, and query results are stored in Amazon S3 with AES-256 server-side encryption, enabled by default on every bucket. No bucket allows public access.
  • The Amazon Aurora PostgreSQL databases have storage encrypted with AWS KMS, which also covers their backups and snapshots.
  • Credentials for source systems are stored in that encrypted database, and the API never returns them.
  • Encryption keys are managed by AWS. Customer-provided keys (BYOK) are not currently offered.

Encryption in transit is described in Security and access control.

What we share and with whom

These are the external services that receive information derived from your data during processing:

RecipientPurposeWhat it receivesRetention
Anthropic and OpenAILanguage models for the data agentsSchemas, per-column statistics, and bounded samples of values (see the detail below)According to the provider’s API terms, which exclude that traffic from training
LangSmith and LangfuseAgent traces, to debug and improve their qualityWhat goes into and comes out of each model call, including the samples14 days
Temporal CloudRun orchestrationExecution metadata: table names, schemas, queries, and the status of each stepDuring the run and a short period afterwards
SentryError monitoringTechnical error informationLimited period

The complete list of providers, including those for identity, billing, and application analytics, is in Providers that process data.

What each data agent receives

AgentWhat it receivesIncludes real values?
Data cleaningSchema, per-column statistics, declared keys, and a sample of up to 100 rowsYes, the sample
Results tablesThe user’s request, the catalog and schemas of the tables, the project knowledge, the most frequent values of each column, and the first rows of the queries it testsYes, frequent values and test rows
Type reconciliationSchemas and statistics of the columns used in joinsNo
DocumentationSchemas, column descriptions, and cleaning queriesNo

Samples are sent unmasked, because the agent needs to see real values to clean well. See How to limit what is shared.

What we do not do

  • We do not train our own models or fine-tune on customer data.
  • We do not sell or share customer data with other customers, or with third parties beyond those listed on this page.
  • We do not write to source systems: we only read.
  • We do not replicate data to other AWS regions.
  • We do not keep the history of the source: the current version of each table is kept, plus 3 days of previous versions.

How to limit what is shared

  • Choose which tables to bring in. Only the tables you select when configuring the source are processed.
  • Exclude columns at the source. If a table contains personal or sensitive data that is not needed for the analysis, we recommend exposing a view without those columns and connecting the view.
  • Use a read-only user with permissions limited to the tables you plan to bring in.
  • Delete what you no longer need. Deleting a source or a project deletes its tables and files. See Retention and deletion.