PhaseoPhaseo
PhaseoPhaseo
Checking statusChecking statusVisit status page
Component-level status is unavailable.

Explore

  • Models
  • Chat
  • Compare
  • Providers
  • Apps
  • Rankings
  • Monitor

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs
  • Methodology

Company

  • Blog
  • Pricing
  • Works With
  • Support
  • Privacy
  • Terms

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

Report:Issue·Support

Spotted a data issue or broken page?Open an issueorcontact support

PhaseoPhaseo
ModelsChatCompareProvidersAppsRankings
ModelsChatCompareProvidersAppsRankings
Sign Up
Add modelChat
Nvidia
Llama 3.1 Nemotron 70B Instruct

Overview

Input modalities
-
Output modalities
-
Providers
1 provider
Input context
128,000
Max output
4,000
Release
Oct 2024
Capabilities
ReasoningWebFine-tune

Pricing

Provider
DeepInfra
DeepInfra
Input
$1.20 / M tokens
Output
$1.20 / M tokens
Cached input
- / M tokens
Plan
standard
Source
Pricing source

Performance

Latency (p50)
-
Throughput (p50)
-
Provider latency
-
Provider throughput
-
Visualize performance
View charts

Activity

30d tokens
0
Total requests
0
Requests in 30m
0

Benchmarks

Shared wins
0
Comparable tests
0
Total results
12
Benchmark charts
View detail

Simulate a response

Estimated input
21 tokens
Context fit
Fits
Estimated cost
$0.0010
Est. response time
-
Pricing basis
$1.20 in / $1.20 out

Overview

Input/output modalities and key model metadata from the catalog.

Nvidia
Llama 3.1 Nemotron 70B Instruct
Nvidia
Available1 priced provider
Input Modalities
-
Output Modalities
-
ReleaseOct 2024
Knowledge CutoffDec 2023
Context128,000
Max Output4,000
LicenseLlama 3.1 Community License

Gateway Usage

30-day activity plus recent runtime. Text-first models use token volume; other modalities fallback to request activity.

Last 30d
Nvidia
Llama 3.1 Nemotron 70B Instruct
Nvidia
0
tokens · last 30 days
Token data up to 25 Jul 2026
Requests
0
Latency
-
Throughput
-
Request activity · 24h0 in 30m

Benchmarks Comparison

Only benchmarks with comparable results across every selected model are shown.

Benchmark Scores (%)

Switch benchmark type to compare percent and numerical families separately.

Nvidia
Llama 3.1 Nemotron 70B Instruct
ARC-C
%Higher is better
Nvidia
Llama 3.1 Nemotron 70B Instruct
69.2%
GSM8K
%Higher is better
Nvidia
Llama 3.1 Nemotron 70B Instruct
91.43%
GSM8K Chat
%Higher is better
Nvidia
Llama 3.1 Nemotron 70B Instruct
81.88%
HellaSwag
%Higher is better
Nvidia
Llama 3.1 Nemotron 70B Instruct
85.58%

Pricing

Per-1M normalized pricing from observed provider tiers. Blended total uses 90% input + 10% output.

Nvidia
Llama 3.1 Nemotron 70B Instruct
Input $/M
$1.20
DeepInfra
Output $/M
$1.20
DeepInfra
Blended $/M
$1.20
90/10 input-output
Pricing
Blended (90/10)Input $/MOutput $/M
Pricing by meter

All unique meters observed across the selected models.

Meter
Llama 3.1 Nemotron 70B Instruct
Best option
Input Text Tokens$1.20
Output Text Tokens$1.20

Availability

API provider availability and subscription plans.

API Availability

Nvidia
Llama 3.1 Nemotron 70B InstructNvidia
Providers
DeepInfra

Subscription Plans