Skip to content
Geocentric
ModelsTechnologySafetyCompanyNewsCareers
Try Arc

Legal

Training Data Transparency

Last updated: 9 September 2026

What this means

California Civil Code section 3111, added by AB 2013, requires a developer that makes a generative AI system publicly available to Californians to post documentation about the data used to train it. This page is that documentation.

Arc was trained on public datasets, not on data we bought. User conversations are a separate source, used only where the user has opted in, and never used to train a model that was already finished before they did.

On this page

Which models this coversArc — 120.8M parametersConversations as a future training sourceUnreleased modelsHow this page is maintained

Which models this covers

Arc is the only Geocentric model publicly available, and the only one requiring disclosure. Two further models are in training and unreleased; when either is made publicly available, its disclosure will be published here before or at release.

Arc — 120.8M parameters

Decoder-only transformer, 120,787,072 parameters, 1,024-token context, 32,000-token vocabulary trained on the corpus below. Released as an experimental checkpoint; see the model card.

Sources and owners of the datasets

DatasetPublished byStageShare
FineWeb-Edu (sample-10BT)HuggingFaceFWPretraining~70% of the pretraining corpus
Cosmopedia v2HuggingFaceTBPretraining~30% of the pretraining corpus
smol-smoltalkHuggingFaceTBInstruction tuningThe full SFT stage

All three are publicly published datasets obtained from their public distribution points. None were purchased, and none were licensed under a negotiated agreement; each is used under the terms its publisher attaches to it.

Pending internal confirmation. Additional corpora — filtered web text, mathematics, scientific abstracts and encyclopedic text — are present in the training environment and may have contributed to this checkpoint. The exact final mixture has not been confirmed against the checkpoint's own training metadata. This item must be resolved and this page corrected before it is treated as a final AB 2013 disclosure.

Purpose of the datasets

FineWeb-Edu supplies educational-quality filtered web text, chosen because it reaches a given quality on far fewer tokens than raw web crawl, which is what makes training on a single consumer GPU feasible. Cosmopedia v2 supplies synthetic textbook and story prose, which teaches clean explanatory style at small scale. smol-smoltalk supplies instruction-following conversations filtered specifically for small models, because full-size chat datasets are too difficult for a model of this size.

Number of data points and types of data

The training recipe targets approximately 3 billion tokens of pretraining text, drawn from millions of documents, followed by an instruction-tuning stage on the order of hundreds of thousands of conversations. All data is natural-language text and code-containing text. There is no image, audio, video, or biometric data: Arc has no vision capability.

Dates of use

EventDate
Datasets collected2026-09
First used in training2026-09-04
Last used in training2026-09-07

The web-derived material reflects the crawl periods of its upstream sources, which predate the datasets' publication and are documented by those publishers.

Copyright, trademark and other protected material

Yes — the pretraining corpus is derived in substantial part from publicly accessible web text, and that text can be expected to include material protected by copyright, and to include trademarks. We did not attempt to assemble a corpus consisting only of public-domain material, and we do not claim to have done so.

Personal information

Yes — web-derived text can contain personal information about individuals, in the ordinary sense that public web pages mention people. We did not deliberately collect personal information, and we did not use any dataset assembled for the purpose of profiling individuals.

Aggregate consumer information

No dataset of aggregate consumer information was used.

Synthetic data

Yes. Cosmopedia v2 is a synthetically generated corpus, and smol-smoltalk is a filtered synthetic instruction dataset. Together they account for a substantial minority of the pretraining corpus and the whole of the instruction-tuning stage.

Cleaning, processing and modification

Documents shorter than 200 bytes were discarded. Documents were concatenated with explicit end-of-document separators so that sequence boundaries fall at real document boundaries. A 32,000-token byte-pair vocabulary was trained on a sample of the corpus, the corpus was tokenized into binary shards, and 0.2% was held out for validation. For the instruction stage, conversations without an assistant turn or with fewer than two messages were dropped, and loss was computed on assistant turns only.

User conversations

No user conversations were used to train the released Arc checkpoint. Training on it finished on 2026-09-07, before the chat service carried meaningful public traffic, and the corpus above is its complete training input.

Conversations as a future training source

Going forward, prompts and model responses from the chat service may be used to train and align Geocentric models — but only from users who have turned on the opt-in setting:

Send my prompts and Geocentric's responses to help improve and train Geocentric models.

Conversations from users who have not turned it on are not retained at all, so there is nothing of theirs that could be included. Where a user has opted in, the stored copy is capped at 30 days, as described in the Privacy Policy, and is never sold.

Where a future model is trained partly on consented conversations, this page will say so, will state the approximate volume and the period the conversations were collected in, and will describe how the data was filtered before use. We will not fold user data into a disclosure that only lists public datasets.

Unreleased models

Two further Geocentric models are in training. Neither is publicly available, so neither requires disclosure under AB 2013 yet. Their disclosures will be prepared from their training manifests and published here when they are released.

How this page is maintained

The content above is derived from our training recipe, dataset download tooling, and data preparation configuration. Where the repository does not establish a fact, we mark it as pending rather than estimate it. If you believe something here is inaccurate, write to contact@geocentricai.com and we will check it against our records.

Models

Model family Arc model card Try Arc

Research

Technology EPICYCLE PARALLAX Training data

Company

About Safety News Careers Contact

Legal

Legal hub Terms of Service Privacy Policy Acceptable Use AI disclosure Security Consent settings
Geocentric © Geocentric Measure first.