A transformer emits one vector per token, but retrieval needs one vector per text. The signal is knowing the pooling options and the catch that almost everyone misses: the model has to be trained for whichever one you pick.
Unlock the other 754 answers · ₹2,000 / $25includes both full courses · progress stays saved · 6 months · one payment · no auto-renew
