Vector indexing all of Wikipedia on a laptop | Better HN

140 comments

103 comments · 21 top-level

traverseda2y ago· 22 in thread

>Disable swap before building the index. Linux will aggressively try to cache the index being constructed to the point of swapping out parts of the JVM heap, which is obviously counterproductive. In my test, building with swap enabled was almost twice as slow as with it off.

This is an indication to me that something has gone very wrong in your code base.

nisa2y ago

> This is an indication to me that something has gone very wrong in your code base.

I'm not sure on what planet all of these people here live that they have success with Linux swap. It's been broken for me forever and the first thing I do is disable it everywhere.

nolist_policy2y ago

Linux swap has been fixed on Chromebooks for years thanks to MGLRU. It's upstream now and you can try it with a recent enough kernel with

  echo y >/sys/kernel/mm/lru_gen/enabled

funcDropShadow2y ago

On a planet where the heap size of the JVM is properly set.

zero_k2y ago

100% agree

[article author]

TBH this was sloppy on my part. I tested multiple runs of the index build and early on kswapd was super busy. I assumed Linux was just caching recently read parts of the source dataset, but it's also possible it was something external to the index build since it's my daily driver machine. After I turned off swap I had no issues and didn't look into it harder.

cwillu2y ago

The usual thing would be to fadvise(POSIX_FADV_DONTNEED) the relevant file handle you don't want cached.

Edit: see for instance https://insights.oetiker.ch/linux/fadvise.html

dekhn2y ago

As a workaround, you can mlock your process which should prevent the application pages from being evicted by swap.

FWIW this is what Cassandra does on startup, but it didn't seem worth it to go to the trouble of dealing with Unsafe for a demo project.

nextaccountic2y ago

Can't mlock be wrapped out in a safe API?

Linux swap has some fuzzy logic that I never fully understood. There have been times I've disabled it because it wasn't doing what I wanted.

lynx232y ago

I almost always reduce swappiness on a new install, the default of (60?) never served me well.

Yea this was strange, I've only ever seen swaps when my main memory was near being full. Maybe they're storing all embeddings in memory?

I routinely see Linux page memory out to swap while having 10+GB free. I can only guess that it really really really likes to cache recently used data from disk.

I've seen the same. It will sometimes prefer to swap rather than evict the disk cache.

I don't know how Linux does this in particular, but intuitively swapping can make sense if part of your allocated RAM isn't being accessed often and the disk is. The kernel isn't going to know for sure of course, and seems in my case it guessed wrong.

All the major web companies disable swap so no work has been done to optimize it.

Yeah, in most situations I'd rather kill processes with excessive mem usage (possibly due to memleak) than have the machine grind to a halt by swapping. Sometimes my headless lab machines would become practically inaccessible over SSH if I didn't swapoff before running something that accidentally chews up mem.

I'll let my personal laptop swap, though. Especially if my wife is also logged in and has tons of idle stuff open.

Could just turn down swappiness?!

AnotherGoodName2y ago

Memory managed languages won’t do long term garbage collection till memory is nearly full.

I suspect they just need to pass in -xmx options to the jvm to avoid this.

threeseed2y ago

Not sure what long term garbage collection means. Because the JVM will GC all objects at any point if they are unused.

If you're referring to full GC you can configure how often that happens and by default it doesn't just wait until memory is nearly full.

marginalia_nu2y ago

Even if this wasn't factually incorrect, default max heap size on the JVM is 25% of system RAM.

menacingly2y ago

I actually wish this were true because of more predictable gc pauses

threeseed2y ago

With G1GC there is a setting called MaxGCPauseMillis which can give you predictability.

gfourfour2y ago· 21 in thread

Maybe I’m missing something but I’ve created vector embeddings for all of English Wikipedia about a dozen times and it costs maybe $10 of compute on Colab, not $5000

bunderbunder2y ago

This is covering 300+ languages, not just English, and it's specifically using Cohere's Embed v3 embeddings, which are provided as a service and currently priced at US$0.10 per million tokens [1]. I assume if you're running on Colab you're using an open model, and possibly a relatively lighter weight one as well?

[1]: https://cohere.com/pricing

traverseda2y ago

This is pretty early in the game to be relying on proprietary embeddings, don't you think? If if they are 20% better, blink and there will be a new normal.

It's insane to me that someone, this early in the gold rush, would be mining in someone else's mine, so to speak

gfourfour2y ago

Ah didn’t realize it was every language. Yes I’m using a light weight open model - but also my use case doesn’t require anything super heavy weight. Wikipedia articles are very feature-dense and differentiable from one another. It doesn’t require a massive feature vector to create meaningful embeddings.

tjakeOP2y ago

it's 35M 1024 vectors Plus the text

janalsncm2y ago

I also don’t quite understand the value of embedding all languages into the same database. If I search for “dog” do I really need to see the same article 300 times?

As a first step they are using PQ anyways. It seems natural to just assume all English docs have the same centroid and search that subspace with hnswlib.

It's split by language. TFA builds an index on the English language subset.

hey, 3 cents cheaper than text-embedding-3-large (without batching)!

Are there some benchmarks available that compare it with the openai model?

Did you do use the same method, i.e. split by chunks each article and vectorize each chunk?

dudus2y ago

That's the only way to do it. You can't index the whole thing. The challenge is chunking. There are several different algorithms to chunk content for vectorization with different pros and cons.

gfourfour2y ago

Yes

janalsncm2y ago

Also, if you’re spending $5000 to compute embeddings, why are you indexing them on a laptop?

syllogistic2y ago

He's not though because cohere stuck the already embedded dataset on huggingface https://huggingface.co/datasets/Cohere/wikipedia-22-12-en-em...

emmelaich2y ago

Got any details?

gfourfour2y ago

Nothing too crazy, just downloading a dump, splitting it into manageable batch sizes, and using a lightweight embedding model to vectorize each article. Using the best GPU available on colab it takes maybe 8 hours if I remember correctly? Vectors can be saved as NPY files and loaded into something like FAISS for fast querying.

This probably deserves its own article and might be of interest to the HN community.

What is the end task(e.g. RAG, or just vector search for question answering), are you satisfied with results in terms of quality?

thomasfromcdnjs2y ago

Did you chunk the articles? If so, in what way?

j0hnyl2y ago

How big is the resulting vector data?

bytearray2y ago

Do you have a link to the notebook?

gfourfour2y ago

No haha just a rats nest of a bunch of notebooks

DataDaemon2y ago

How?

Atotalnoob2y ago· 5 in thread

Why is the author listing himself as datastax cto?

He isn’t according the Wikipedia, my friend who works there, and their company website. https://www.datastax.com/our-people

That’s kind of weird

I guess I'm kind of a CTO emeritus now -- I mostly write code, by choice. https://github.com/jbellis

See https://www.datastax.com/our-people/jonathan-ellis

Wikipedia lists them as a founder. Perhaps their author bio is outdated, or Wikipedia is. Not sure about your friend.

Atotalnoob2y ago

They were definitely a founder, but they are not the current cto

What are you talking about? The datastax site lists it:

> SANTA CLARA, Calif. – September 28, 2020 – DataStax today announced that DataStax Co-Founder and CTO Jonathan Ellis will deliver a keynote address at ApacheCon @Home 2020

https://www.datastax.com/press-release/datastax-co-founder-a....

As an aside, I'm an ApacheCon presenter but there was no press release about the hot excitement of my involvement. Maybe next time :)

Atotalnoob2y ago

That’s from 2024. They aren’t the cto of datastax currently

tjakeOP2y ago· 4 in thread

You can demo this here: https://jvectordemo.com:8443/

GH Project: https://github.com/jbellis/jvector

[article author]

The source to build and serve the index are at https://github.com/jbellis/coherepedia-jvector

Internal sever error :(

Oops, HN maxed out the free Cohere API key it was using. Fixed.

bytearray2y ago

Same.

HammadB2y ago· 4 in thread

"The obstacle is that until now, off-the-shelf vector databases could not index a dataset larger than memory, because both the full-resolution vectors and the index (edge list) needed to be kept in memory during index construction. Larger datasets could be split into segments, but this means that at query time they need to search each segment separately, then combine the results, turning an O(log N) search per segment into O(N) overall."

How is a log N search over S segments O(N)?

I was trying to make the point that the dominant factor becomes linear instead of logarithmic, but more accurately it's O(S log N) = O(N log N) because S is proportional to N.

Sure, I see. I think this is an area where complexity analysis doesn’t lead to useful information.

To be more correct it’s O(N/C log C) where C is the capacity of a segment. In this case you can ignore 1/C and log C as constant. So now sure, you actually just have O(N). But this is not super useful as it says that a segmented hnsw approach and brute force approach are the same - when this is really not the case in practice.

Also O(N log N) > O(N) so I’m not sure why we would ever do anything with segmentation according to that analysis if it were correct.

> I’m not sure why we would ever do anything with segmentation according to that analysis if it were correct.

What's your alternative when you can't build an index larger than C?

Doesn't doubling N double S?

noufalibrahim2y ago· 4 in thread

What are the good solutions in this space? Vector databases I mean. Mostly for semantic search across various texts.

I have a few projects I'd like to work on. For typical web projects, I have a "go to" stack and I'd like to add something sensible for vector based search to that.

In my experience its usually easiest to use a vector store extension for an off-the-shelf database like postgres (pgvector is nice). That way you don't have to manage another, rapidly changing, service and you can easily combine queries on the vectors with regular columns, join them and so on.

JVector (the index used in TFA) is available as a service with a friendly API from DataStax. https://www.datastax.com/products/datastax-astra

[article author, I work on JVector and Astra]

Could you tell how scalable JVector is? How many vectors it can handle, like millions, billions, hundreds of billions?

noufalibrahim2y ago

Nice. I wanted to try something out on a machine before moving to hosted soclutions.

m3kw92y ago· 4 in thread

He should have asked HN on the cheapest way to embed Wikipedia before starting

I'm baffled that so many people fixate on the estimated cost and miss the fact that it's a public dataset. As in, free.

syllogistic2y ago

It's in the first sentence of the article too :)

m3kw92y ago

Getting the embeddings ain’t free

Unless you're counting your network access cost, the dataset of embeddings is in fact free and TFA includes instructions on how to download them for free.

isoprophlex2y ago· 3 in thread

$5000?! I indexed all of HN for ... $50 I think. And that's tens of millions of posts.

To be fair Wikipedia has over 60 million pages and this is for 300+ languages. But yeah, the value shows that they might not be using the cheapest service out there.

How are you using that index?

isoprophlex2y ago

https://www.searchhacker.news/

A tool that (hopefully) surfaces interesting HN discussion threads; I wanted an excuse to investigate (hybrid) full text and vector search at a substantial scale beyond toy datasets.

Sadly (well not really) I changed jobs soon after building the first version. Life caught up and I never got around to adding more features and polishing up the frontend (eg. the broken back button

Ideas for new features are very welcome :)

hot_gril2y ago· 3 in thread

How many dimensions are in the original vectors? Something in the millions?

1024 per vector x 41M vectors

1024-dim vectors would fit into pgvector in Postgres, which can do cosine similarity indexing and doesn't require everything to fit into memory. Wonder how the performance of that would compare to this.

It's been a while since I read the source to pgvector but at the time it was a straightforward HNSW implementation that implicitly assumes your index fits in page cache. Once that's not true, your search performance will fall off a cliff. Which in turn means your insert performance also hits a wall since each new vector requires a search.

I haven't seen any news that indicates this has changed, but by all means give it a try!

lilatree2y ago· 3 in thread

“… turning an O(log N) search per segment into O(N) overall.”

Can someone explain why?

StrangeDoctor2y ago

When it’s all in memory you get to amortize the cost of the initial load. Or just pay it when it’s not part of the hot path. When it’s segmented, you’re doing that because memory is full and you need to read in all the segments you don’t have. That’ll completely overwhelm the log n of the search you still get

I was trying to make the point that the dominant factor becomes linear instead of logarithmic, but more accurately it's O(S log N) = O(N log N) because S (number of segments) is proportional to N (number of vectors).

StrangeDoctor2y ago

Ah yeah that’s what I wanted to write but I guess I didn’t want to put words in your mouth, and stuck to what I could be certain about happening. We do all this work to throw away the unneeded bits in one situation and when comparing it to a slightly different situation go “huh some of that garbage would be kinda nice here”

peter_l_downs2y ago· 2 in thread

> JVector, the library that powers DataStax Astra vector search, now supports indexing larger-than-memory datasets by performing construction-related searches with compressed vectors. This means that the edge lists need to fit in memory, but the uncompressed vectors do not, which gives us enough headroom to index Wikipedia-en on a laptop.

It's interesting to note that JVector accomplishes this differently than how DiskANN described doing it. My understanding (based on the links below, but I didn't read the full diff in #244) is that JVector will incrementally compress the vectors it is using to construct the index; whereas DiskANN described partitioning the vectors into subsets small enough that indexes can be built in-memory using uncompressed vectors, building those indexes independently, and then merging the results into one larger index.

OP, have you done any quality comparisons between an index built with JVector using the PQ approach (small RAM machine) vs. an index built with JVector using the raw vectors during construction (big RAM machine)? I'd be curious to understand what this technique's impact is on the final search results.

I'd also be interested to know if any other vector stores support building indexes in limited memory using the partition-then-merge approach described by DiskANN.

Finally, it's been a while since I looked at this stuff, so if I mis-wrote or mis-understood please correct me!

- DiskANN: https://dl.acm.org/doi/10.5555/3454287.3455520

- Anisotropic Vector Quantization (PQ Compression): https://arxiv.org/abs/1908.10396

- JVector/#168: How to support building larger-than-memory indexes https://github.com/jbellis/jvector/issues/168

- JVector/#244: Build indexes using compressed vectors https://github.com/jbellis/jvector/pull/244

It's not mentioned in the original paper, but DiskANN also supports PQ at build-time via `--build_PQ_bytes`, though it's a tradeoff with the graph quality as you mention.

One interesting property in benchmarking is that the distance comparison implementations for full-dim vectors can often be more efficient than those for PQ-compressed vectors (straight-line SIMD execution vs table lookups), so on some systems cluster-and-merge is relatively competitive in terms of build performance.

That's correct!

I've tested the build-with-compression approach used here with all the datasets in JVector's Bench [1] and there's near zero loss in accuracy.

I suspect that the reason the DiskANN authors used the approach they did is that in 2019 Deep1B was about the only very large public dataset around, and since the vectors themselves are small your edge lists end up dominating your memory usage. So they came up with a clever solution, at the cost of making construction 2.5x as expensive. (Educated guess: 2x is from adding each vector to multiple partitions and the extra 50% to merge the results.)

So JVector is just keeping edge lists in memory today. When that becomes a bottleneck we may need to do something similar to DiskANN but I'm hoping we can do better because it's frankly a little inelegant.

[1] https://github.com/jbellis/jvector/blob/main/jvector-example...

jl62y ago· 2 in thread

The source files appear to include pages from all namespaces, which is good, because a lot of the value of Wikipedia articles is held in the talk page discussions, and these sometimes get stripped from projects that use Wikipedia dumps.

worldsayshi2y ago

I'm curious what the main value you see in the talk pages? I almost never look at them myself.

jl62y ago

They’re not so interesting for mundane topics, but for anything remotely controversial, they are essential for understanding what perspectives aren’t included in the article.

Mathnerd3142y ago· 2 in thread

> Enough RAM to run a JVM with 36GB of heap space

Are there laptops like that? Maybe an upgraded MacBook, but I have been looking for Windows/Linux laptops and they generally top out at 32GB. I checked Lenovo's website and everything with 64GB and up is not called a laptop but a "mobile workstation".

You can configure a Lenovo Z13 Gen 2 with 64GB for little extra money (and choose between Windows, Ubuntu, Fedora, or no OS preinstalled).

You can buy an M3 Max with 128GB memory.

issafram2y ago· 2 in thread

Would a docker container help running it on Windows?

Technically it does run on windows, you just can't build the entire dataset without adding the sharding code mentioned. Set divisor=100 in config.properties and it will happily build an index over 1% of the dataset.

Just use WSL. Or dual boot.

arnaudsm2y ago· 1 in thread

In expert topics, is vector search finally competitive with BM25-like algorithms? Or do we still need to mix the 2 together ?

ColBERT gives you best of both worlds.

https://arxiv.org/abs/2004.12832

https://thenewstack.io/overcoming-the-limits-of-rag-with-col...

burgerrito2y ago

I made a side project that uses Wikipedia recently too, and found out that there are database dump available to be downloaded: https://en.wikipedia.org/wiki/Wikipedia:Database_download

localhost2y ago

This is a giant dataset of 536GB of embeddings. I wonder how much compression is possible by training or fine-tuning a transformer model directly using these embeddings, i.e., no tokenization/decoding steps? Could a 7B or 14B model "memorize" Wikipedia?

anonymousDan2y ago

How do embeddings created by state of the art open source models compare to the free embeddings mentioned in the article? Would they actually cost 5k to create given a reasonable local GPU setup?

opdahl2y ago

Would be interesting if you could try implementing the Cohere Reranker into this. Should be fairly easy, and could lead to quite a bit of performance gain.

Khelavaster2y ago

This is how Microsoft powered it's academic paper search in 2016, before rolling it into Bing in 2020!

Loosely related: https://www.quantamagazine.org/computer-scientists-invent-an...

j / k navigate · click thread line to collapse