Researchers Translate Embeddings Between AI Models Without Any Paired Data

Harnessing the Universal Geometry of Embeddings

A new paper from Cornell Tech introduces the first method to translate text embeddings between different vector spaces without paired data, encoders, or predefined matches. The unsupervised approach maps embeddings to and from a universal latent representation, achieving high cosine similarity across models with different architectures, parameter counts, and training data. This capability has serious security implications: an adversary with only embedding vectors could extract sensitive information from underlying documents, enabling classification and attribute inference.

An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.
  1. srean

    One way to pose/(think about) the problem is that there are two finite metric spaces linked by an unknown odometry (damn you autocorrect). The problem is to recover that unknown isometry.

    This, like graph isometry, can be very computationally intensive in the worst case. However, heuristics to aid matching one vertex on one graph to another vertex on another graph using local, semilocal structural signatures can be very effective on particular cases.

    One can of course argue that the spaces are not designed as metric spaces. Even if true, these might be metrizable topological spaces.

    More generally, if these are indeed non-metric spaces one can still pose it as finding the unknown isomorphism between two poset spaces.

    In my other comment I was using the property of maximal chains -- Identify the longest chains in both posets. The isomorphism must map the longest chain in Poset 1 directly to a longest chain in Poset 2, preserving the exact linear order.

  2. nickledave

    Dupe: https://news.ycombinator.com/item?id=44054425

    Note this is version 4 of the paper and the original post was version 1 (I think?)

    OpenReview (for NeurIPS) for the curious: https://openreview.net/forum?id=jiCLUPq5xv

  3. ironSkillet

    I am not familiar with the standards of publishing in machine learning, but as someone trained in a mathematics background, this paper seems relatively light on details and heavy on exposition. Is that typical? Is this a really novel idea? Not trying to be snarky, just trying to understand how meaningful this is.

  4. fennecfoxy

    I'm not as heavy on the maths stuff involved in this as other people commenting appear to be.

    But the idea makes sense, of course there is still recoverable data in embeddings, that's the point. Though as I constantly find the more you try to squeeze into an n bit vector the more watered down everything gets.

    I suppose a latent space could be encrypted/mapped in some way to resolve that, but how many people are exposing their vectors in the first place?

  5. srean

    Let's assume that monotonocity of pair-wise distances are preserved.

    Without knowing the details of how the paper solved the problem, my first attempt would be to find the diametrically distant pair of points in the two different embeddings and assume that the pair is the same pair. Then find the next distant pairs and so on.

    After sufficiently many such pairs have been found, or better still, the largest d-simplex is found, find that scaled rigid body transformation that makes the corresponding pairs coincide. Proceeding this way ought to be less work than solving a generic graph isomorphism problem.

More from this day

2026-09-06