Skip to content

All Content tagged with AWS Inferentia

AWS Inferentia is designed to provide high performance inference in the cloud, to drive down the total cost of inference, and to make it easy for developers to integrate machine learning into their business applications.

Content language: English

Filter content
Select tags to filter
Sort by
Sort by most recent
57 results
New Training & Certification Badge from AWS
Virtual training on choosing the optimal infrastructure for Small Language Models
[AWS Neuron Documentation](https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/setup/neuron-setup/multiframework/multi-framework-ubuntu22-neuron-dlami.html#setup-ubuntu22-multi-framework-d...
1
answers
0
votes
185
views
AWS

asked 2 years ago

Walk through the options for compiling a model for inference using Inferentia or Trainium. You would need to do this if the model or the configuration you want isn't available in the Hugging Face cac...
AWS

KamranEXPERT

published 2 years ago3 votes5.5K views

Step by Step guide to deploy DeepSeek R1 Distilled models.
AWS

published 2 years ago0 votes634 views

A list of resources to use when you are first starting with the Neuron SDK and Inferentia or Trainium instances.
Steps to set up Jupyter notebooks and VS Code remote server on Trainium and Inferentia Neuron systems.
AWS

KamranEXPERT

published 2 years ago0 votes1.7K views

Key announcements and discover how industry leaders, like Apple and Anthropic, are revolutionizing AI with AWS Trainium and Inferentia
Get started with Inferentia and Trainium on EC2 using the Hugging Face Neuron Deep Learning Amazon Machine Image (AMI). A short walkthrough of how to deploy an EC2 image with all the Neuron drivers a...
AWS

KamranEXPERT

published 2 years ago0 votes1.3K views

Are you heading to **AWS re:Invent 2024** and looking for AWS Inferentia and Trainium sessions to take your machine learning skills to the next level?
See what regions have instances, and find out how to generate your own list with a python script.
Hello AWS team! I am trying to run a suite of inference recommendation jobs leveraging NVIDIA Triton Inference Server on a set of GPU instances (ml.g5.12xlarge, ml.g5.8xlarge, ml.g5.16xlarge) as well...
1
answers
0
votes
1K
views

asked 2 years ago

  • 1
  • 2
  • 3
  • 4
  • 5
  • Page size
    12 / page