$ cat publications.md
Selected publications below — full list on Google Scholar.
In this report, we introduce the Gemini 2.X model family:
Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and
Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving
SoTA performance on frontier coding and reasoning benchmarks. In addition
to its incredible coding and reasoning skills, Gemini 2.5 Pro is a
thinking model that excels at multimodal understanding and it is now able
to process up to 3 hours of video content. Its unique combination of long
context, multimodal and reasoning capabilities can be combined to unlock
new agentic workflows. Gemini 2.5 Flash provides excellent reasoning
abilities at a fraction of the compute and latency requirements and
Gemini 2.0 Flash and Flash-Lite provide high performance at low latency
and cost. Taken together, the Gemini 2.X model generation spans the full
Pareto frontier of model capability vs cost, allowing users to explore
the boundaries of what is possible with complex agentic problem solving.
Sparse mixture of expert architectures (MoEs) scale model capacity
without significant increases in training or inference costs. Despite
their success, MoEs suffer from a number of issues: training instability,
token dropping, inability to scale the number of experts, or ineffective
finetuning. In this work, we propose Soft MoE, a
fully-differentiable sparse Transformer that addresses these challenges, while maintaining the
benefits of MoEs. Soft MoE performs an implicit soft assignment by
passing different weighted combinations of all input tokens to each
expert. As in other MoEs, experts in Soft MoE only process a subset of
the (combined) tokens, enabling larger model capacity (and performance)
at lower inference cost. In the context of visual recognition, Soft MoE
greatly outperforms dense Transformers (ViTs) and popular MoEs (Tokens
Choice and Experts Choice). Furthermore, Soft MoE scales well: Soft MoE
Huge/14 with 128 experts in 16 MoE layers has over 40x more parameters
than ViT Huge/14, with only 2% increased inference time, and
substantially better quality.
Training large, deep neural networks to convergence can be prohibitively
expensive. As a result, often only a small selection of popular, dense
models are reused across different contexts and tasks. Increasingly,
sparsely activated models, which seek to decouple model size from
computation costs, are becoming an attractive alternative to dense
models. Although more efficient in terms of quality and computation
cost, sparse models remain data-hungry and costly to train from scratch
in the large scale regime. In this work, we propose
sparse upcycling – a simple way to reuse
sunk training costs by initializing a sparsely activated
Mixture-of-Experts model from a dense checkpoint. We show that
sparsely upcycled T5 Base, Large, and XL language models and Vision
Transformer Base and Large models, respectively, significantly
outperform their dense counterparts on SuperGLUE and ImageNet, using
only ~50% of the initial dense pretraining sunk cost. The upcycled
models also outperform sparse models trained from scratch on 100% of
the initial dense pretraining computation budget.
Effective scaling and a flexible task interface enable large language
models to excel at many tasks. We present PaLI (Pathways Language and
Image model), a model that extends this approach to the joint modeling
of language and vision. PaLI generates text based on visual and textual
inputs, and with this interface performs many vision, language, and
multimodal tasks, in many languages. To train PaLI, we make use of large
pre-trained encoder-decoder language models and Vision Transformers
(ViTs). This allows us to capitalize on their existing capabilities and
leverage the substantial cost of training them. We find that
joint scaling of the vision and language components is important. Since
existing Transformers for language are much larger than their vision
counterparts, we train a large, 4-billion parameter ViT (ViT-e) to
quantify the benefits from even larger-capacity vision models. To train
PaLI, we create a large multilingual mix of pretraining tasks, based on a
new image-text training set containing 10B images and texts in over 100
languages. PaLI achieves state-of-the-art in multiple vision and
language tasks (such as captioning, visual question-answering, scene-text
understanding), while retaining a simple, modular, and scalable design.
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated
excellent scalability in Natural Language Processing. In Computer
Vision, however, almost all performant networks are “dense”, that is,
every input is processed by every parameter. We present a
Vision MoE (V-MoE), a
sparse version of the Vision Transformer, that is scalable
and competitive with the largest dense networks. When applied to image
recognition, V-MoE matches the performance of state-of-the-art networks,
while requiring as little as half of the compute at inference time.
Further, we propose an extension to the routing algorithm that can
prioritize subsets of each input across the entire batch, leading to
adaptive per-image compute. This allows V-MoE to trade-off performance
and compute smoothly at test-time. Finally, we demonstrate the potential
of V-MoE to scale vision models, and train a 15B parameter model that
attains 90.35% on ImageNet.
Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning
when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised
datasets and fine-tuning the model on a target task. We scale up pre-training, and propose a simple recipe
that we call Big Transfer (BiT). By combining a few carefully selected components, and transferring using
a simple heuristic, we achieve strong performance on over 20 datasets. BiT performs well across a
surprisingly wide range of data regimes – from 1 example per class to 1M total examples.
BiT achieves 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19 task
Visual Task Adaptation Benchmark (VTAB). On small datasets, BiT attains 76.8% on ILSVRC-2012 with 10
examples per class, and 97.0% on CIFAR-10 with 10 examples per class. We conduct detailed analysis of
the main components that lead to high transfer performance.
Keyword Spotting, applied to handwritten text documents, aims to retrieve
the documents, or parts of them, that are relevant for a query, given by
the user, within a large collection of documents. The topic has gained a
large interest in the last 20 years among Pattern Recognition
researchers, as well as digital libraries and archives. This thesis,
first defines the goal of Keyword Spotting from a Decision Theory
perspective. Then, the problem is tackled following a
probabilistic formulation. More precisely, Keyword Spotting is presented as a
particular instance of Information Retrieval, where the content of the
documents is unknown, but can be modeled by a probability distribution.
In addition, the thesis also proves that, under the correct probability
distributions, the framework provides the optimal solution, under many
of the evaluation measures traditionally used in the field. Later,
different statistical models are used to represent the probability
distribution over the content of the documents. These models, Hidden
Markov Models or Recurrent Neural Networks, are estimated from training
data, and the corresponding distributions over the transcripts of the
images can be efficiently represented using
Weighted Finite State Transducers. In order to make the framework practical for large
collections of documents, this thesis presents several algorithms to
build probabilistic word indexes, using both lexicon-based and
lexicon-free models. Finally, all the contributions are evaluated
experimentally, not only on standard academic benchmarks, but also on
collections including tens of thousands of pages of historical
manuscripts.
Current state-of-the-art approaches to offline Handwritten Text
Recognition extensively rely on
Multidimensional Long Short-Term Memory networks.
However, these architectures come with quite an expensive computational
cost, and we observe that they extract features visually similar to
those of convolutional layers, which are computationally cheaper.
This suggests that the two-dimensional long-term dependencies, which
are potentially modeled by multidimensional recurrent layers, may not
be essential to achieve a good recognition accuracy, at least in the
lower layers of the architecture. In this work, an alternative model
is explored that relies only on
convolutional and one-dimensional recurrent layers that achieves better or equivalent results than those
of the current state-of-the-art architecture, and runs significantly
faster.
In addition, we observe that using random distortions during
training as synthetic data augmentation dramatically improves the
accuracy of our model.
Thus, are multidimensional recurrent layers really necessary for
Handwritten Text Recognition? Probably not.