Research

Papers & preprints

Original papers and technical reports from our work on GIDE — alongside the foundational third-party research it builds on.

What these topics mean
Distributed Training
Sharding computation and memory across many devices to scale beyond a single accelerator.
Efficient Attention
Methods that cut attention's compute and memory cost without changing its output.
GPU Kernels & Systems
Hardware-aware implementations that map attention efficiently onto GPU memory and compute.
Long-Context
Extending models to very long sequences — hundreds of thousands to billions of tokens.
Low-Precision & Quantization
Trading numerical precision for speed and memory while bounding the accuracy cost.
Safety & Control
Methods that keep a learning system provably within specified safety limits while it acts.
Self-Supervised Learning
Learning useful representations from unlabeled data by predicting parts of the input from other parts.
State-Space Models
Linear-time sequence models that replace attention with a recurrent state, scaling to long sequences without the quadratic cost.
Transformer Architecture
The attention-based neural network architecture that underlies modern sequence models.
World Models
Models that learn how a system evolves in a latent representation space, for prediction and planning rather than generation.

Topic

Transformer Architecture

The attention-based neural network architecture that underlies modern sequence models.

Showing 1 of 18 papers.

Foundational reading

Notable third-party research we build on.

  • Attention Is All You Need

    Published Third-party DarcStar commentary NeurIPS·June 12, 2017

    By Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin

    This is third-party work — not authored by or affiliated with DarcStar Technologies.

    Introduces the Transformer, a sequence-transduction architecture based solely on attention mechanisms, dispensing with recurrence and convolutions.