Category

AI Compute Economics

These terms describe request latency, token output, reuse of cached work, and the memory footprint of a model.

How to recognize this theme

Words that describe how models trade speed for cost and memory.

In a daily board, this category groups terms by their shared role. Look for four cards that describe the same mechanism, risk area, or workflow rather than four words that merely sound similar.

Educational context

These entries are vocabulary notes for learning. They are not project endorsements, token recommendations, exchange rankings, or trading signals.

Inference Speed

Inference speed describes how quickly a model can produce an answer after receiving a request.

Token Rate

Token rate measures how many output tokens a serving system can produce in a given amount of time.

Cache Hit Rate

Cache hit rate is the share of requests that can reuse stored work instead of recomputing it.

Model Footprint

Model footprint is the memory or storage size a model needs while being deployed or served.