Welcome to emu-mps¶
You have found the documentation for emu-mps. The emulator emu-mps is a backend for the Pulser low-level Quantum Programming toolkit that lets you run Quantum algorithms on a simulated device, using GPU acceleration if available. More in depth, emu-mps is designed to emulate the dynamics of programmable arrays of neutral atoms, with matrix product states (mps). While benchmarking is incomplete as of this writing, early results suggest that this design makes emu-mps faster and more memory-efficient than previous generations of quantum emulators at running simulations with large numbers of qubits.
Supported features¶
The following features are currently supported:
- All Pulser sequences that use only rydberg (
ground-rydbergbasis) and only microwave (XYbasis) channel - MPS and MPO can be constructed using the abstract Pulser format and following the correspondent basis format
- All noise from the pulser
NoiseModel- Effective noise (
eff_noise) is included using jump or collapse operators and emu-mps uses the quantum jump method or Monte Carlo wave function (MCWF) approach.
- Effective noise (
- The following basis states in a sequence:
- The following properties from a Pulser Sequence are also correctly applied:
- hardware modulation
- SLM mask
- A complex phase for the omega parameter, i.e. the phase \(\phi\) in the driving Hamiltonian
- Customizable output, with the folowing inbuilt options:
- The quantum state in MPS format
- Bitstrings
- The fidelity with respect to a given state
- The expectation of a given operator (as
MPOorMPO._from_operator_repr) - The qubit density (magnetization)
- The correlation matrix
- The mean, second moment and variance of the energy
- Entanglement entropy
- computational statistics: each time step during the simulation will generate the following information:
- \(\chi\) : is the maximum bond dimension of the MPS
- \(|\Psi|\): MPS (the state) memory footprint
- RSS: max memory allocation. If run on GPU this is cited both for RAM and GPU.
- \(\triangle t\): time that the step took to run (given in seconds)
- Specification of:
- Initial state ( as
MPSorMPS._from_state_amplitudes) - Various precision parameters
- Whether to run on CPU or GPU
- The interaction coefficients \(U_{ij}\) from here
- A cutoff below which \(U_{ij}\) are set to 0 (this makes the computation more memory efficient)
- Initial state ( as
Planned features¶
- Differentiability.
Environment variables¶
The workflow in emu-mps is not typical of machine learning workloads, and as a consequence certain default behaviours of torch are not optimal. Specifically
- In emu-mps tensor sizes tend to be very unpredictable, in contrast to typical machine learning models. The default torch behaviour of caching GPU memory allocations in deterministic block sizes must be overridden to avoid memory fragmentation that reduces the effectively available memory on the GPU.
- In emu-mps, we use page-locked RAM for asynchronously transferring tensors from RAM to GPU. Since the tensor sizes are unpredictable, the default torch caching behaviour of allocating memory in powers of 2 works poorly. It is not needed for performance, and it causes the application to use on average 1.5 times more RAM than needed. The host allocator cache should be disabled. This is only possible in torch 2.13 and newer.
Both these things can be configured via an environment variable that should be set before running your python script. Concretely, define the following for optimal memory usage:
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,pinned_max_cached_size_mb:0
On torch versions lower than 2.13 only use everything before the comma.
More Info¶
Please see the API specification for a list of available config options (see here). Those configuration options relating to the mathematical functioning of the backend are explained in more detail in the config page (see here). For notebooks with examples for how to do various things, please see the notebooks page (see here).