Borderless lakehouse
Query data that lives in several clouds as if it were in one — without copying it anywhere. Pick the cloud that will hold the lakehouse, add one source per cloud your data sits in, tick the tables you want, and you get the price of querying across each boundary plus everything needed to stand it up.
What changes, and what does not
The data does not move. Your tables stay in the buckets and the clouds they are in today. What becomes one thing is the catalogue and the place you query from. That is the whole idea, and it is why the cost below is a cost of reading across a boundary rather than a cost of moving.
This is an example. Add your sources below and the picture redraws from your estate — your clouds, your table counts, your egress.
Two different things are called a “catalogue” — you need both
One knows what a table means. The other knows which files a table is. Mixing them up is the classic mistake on this kind of build, so this page keeps them in separate sections all the way down.
Semantic catalogue — what the data means
Names, owners, business terms, lineage. Without it you have a list of tables nobody can explain: you can see c_stat_01 but not what it holds, who owns it, or who to ask before moving it.
- Google Knowledge Catalog — read for a source on Google Cloud. It is Google's catalogue product, renamed from Dataplex Universal Catalog in April 2026.
- AWS Glue Data Catalog — read for a source already on AWS, because that is where an AWS estate's meaning already lives. Nobody has to re-enter it.
Two adapters on purpose: one catalogue would have let this platform quietly assume every estate is Google-shaped. Glue proves it does not.
Glue is where meaning is read from for an AWS source. It is not an alternative place to put the lakehouse — that is the target you choose in step 1.
Technical catalogue — where the bytes are
The pointer to each Iceberg table's current files and manifests. Without it a query engine cannot find which files a table is made of right now, so nothing is queryable at all — however well documented it is.
- BigLake Metastore — the technical catalogue for a Google Cloud lakehouse. It is a component of the target, so the target you pick decides which one you get.
It holds nothing about what a table means or who owns it. That is the other card.
You do not configure this one. It is built from the tables you tick, and appears at the bottom of this page once you form the lakehouse.
Semantic catalogue — state unknown
The catalogue's state could not be read from this API. That is unknown, not “not configured” — do not conclude either way from this card.
Choose where the lakehouse lives
One cloud holds the catalogue and the query surface. Your data stays where it is — this is not a destination for the bytes.
The target list could not be read from this API. That is unknown, not “no targets”.
A cloud marked not built yet is a gap in Pathfinder, not a limit of that cloud.
Add one source per cloud your data sits in
Point each source at its estate JSON — a collector export, or one of the samples. The cloud you pick sets the price of reading across that boundary and how its tables reach the lakehouse.
Nothing added yet. Add one source per cloud the data lives in — with only one source there is no border to cross, and this is just a lakehouse.
Assess first to see the cost and pick the tables. Form second to get the build.
Both buttons are off because Pathfinder cannot build into the target you picked. Its card in step 1 says what is missing. That is a gap here, not a limit of that cloud.
Choose how it gets built
Then press “Form the lakehouse” in step 2. You can change this and form it again — nothing here is a saved setting.
Same lakehouse, three ways to get it standing up. Pick one before you form it.
These three came from the UI, not from this API. The backend that serves the apply modes is not merged yet, so only the Terraform download really works here. Treat the other two as a preview of the screen, not as a claim about this deployment.