Model evaluations
Score a trained detection or instance-segmentation model against an export's ground truth - precision, recall, F1, mAP, confusion matrix, worst images.
An evaluation runs your trained model over an export and scores its predictions against that export’s ground truth. The inference runs server-side, so you do not download weights or images to get metrics, and your annotations are never modified.
The eval set is the export, not the whole dataset. An export records exactly which images and which classes it contains, so evaluating against one gives you a fixed, repeatable scoring set and puts evaluation on the same footing as training, which also runs off exports.
Only detection and instance-segmentation models can be evaluated. Other
model types return a validation error. For scoring predictions you already hold,
without spending compute, use the offline
pictograph.metrics
helpers - the metric math is identical, so a server run and a local run on the same
data agree.
Evaluation is billed on the trained-model batch-inference path, in compute credits.
There is no pictograph CLI group for evaluations; use the Python SDK or raw REST.
evaluate
The one-call path: start the run and block until it finishes. Use this unless you need to do something else while it runs.
Source: ModelEvaluations.evaluate
| Arg | Type | Default | Notes |
|---|---|---|---|
model |
str |
required | Trained model name, or a model id |
dataset_name |
str |
required | Dataset name. |
export_name |
str |
required | Completed export whose images and class filter are the eval set. Export names are unique within a dataset, not globally, which is why the pair is needed |
iou_threshold |
float |
0.5 |
Minimum IoU for a prediction to count as matching a ground-truth box |
confidence_threshold |
float |
0.5 |
Minimum detection confidence to keep a prediction |
poll_interval |
float |
3.0 |
Seconds between status checks |
timeout |
float |
1800.0 |
Seconds before PollTimeoutError |
evaluation = client.model_evaluations.evaluate(
model="Swift Falcon",
dataset_name="road-signs",
export_name="v1",
iou_threshold=0.5,
confidence_threshold=0.5,
)
overall = evaluation.overall_metrics
print(overall.precision, overall.recall, overall.f1, overall.map)
for row in evaluation.per_class_metrics or []:
print(row.class_name, row.precision, row.recall, row.support, row.ap)
# REST has no combined call: create, then poll get.
curl -s -X POST "https://api.pictograph.io/api/v1/developer/model-evaluations" \
-H "X-API-Key: $PICTOGRAPH_API_KEY" -H "Content-Type: application/json" \
-d '{"model_id": "<model-uuid>", "export_id": "<export-uuid>"}'
Returns ModelEvaluation
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
Raises ApiError if the run failed or was cancelled, PollTimeoutError if timeout elapsed first (the run keeps going, so fetch it later with get).
create
Start an evaluation and return immediately with status pending.
Source: ModelEvaluations.create
| Arg | Type | Default | Notes |
|---|---|---|---|
model |
str |
required | Trained model name, or a model id |
dataset_name |
str |
required | Dataset name. |
export_name |
str |
required | Completed export defining the eval set |
iou_threshold |
float |
0.5 |
Match threshold |
confidence_threshold |
float |
0.5 |
Prediction cutoff |
evaluation = client.model_evaluations.create(
model="Swift Falcon",
dataset_name="road-signs",
export_name="v1",
)
print(evaluation.id, evaluation.status)
curl -s -X POST "https://api.pictograph.io/api/v1/developer/model-evaluations" \
-H "X-API-Key: $PICTOGRAPH_API_KEY" -H "Content-Type: application/json" \
-d '{
"model_id": "6f1c2f0e-6a1e-4a55-9f3a-2f2d3b4c5d6e",
"export_id": "b4d1c7e2-3f58-4a90-8c6b-1d2e3f4a5b6c",
"iou_threshold": 0.5,
"confidence_threshold": 0.5
}'
Returns ModelEvaluation
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
The SDK resolves the model name and the dataset-plus-export pair into ids for you;
REST takes the ids directly. The export must be completed and must carry a
recorded image set, so exports created before per-export image tracking must be
re-created before they can be scored.
get
Fetch one evaluation, including its metrics once it has completed.
Source: ModelEvaluations.get
| Arg | Type | Default | Notes |
|---|---|---|---|
evaluation_id |
str |
required | Evaluation id from create or list |
evaluation = client.model_evaluations.get(
evaluation_id="9d8b3c17-2f45-4e6a-b810-5c7e2a94f036",
)
print(evaluation.status, evaluation.progress, evaluation.evaluated_images)
curl -s "https://api.pictograph.io/api/v1/developer/model-evaluations/9d8b3c17-2f45-4e6a-b810-5c7e2a94f036" \
-H "X-API-Key: $PICTOGRAPH_API_KEY"
overall_metrics carries micro-averaged tp / fp / fn, precision, recall,
f1, macro_f1 and map. Each entry in per_class_metrics adds class_name,
support and per-class ap. confusion_matrix gives labels and a dense grid
where rows are ground truth and columns are predictions; the final label is
__background__, so the last column holds false negatives and the last row holds
false positives. worst_images lists the images with the most errors, each with
filename, tp, fp, fn, gt_count and pred_count.
matrix = evaluation.confusion_matrix
print(matrix.labels)
for label, row in zip(matrix.labels, matrix.grid):
print(label, row)
for image in evaluation.worst_images or []:
print(image.filename, image.fp, image.fn)
Returns ModelEvaluation
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
wait_for_completion
Poll an evaluation you started earlier until it reaches a terminal status.
Source: ModelEvaluations.wait_for_completion
| Arg | Type | Default | Notes |
|---|---|---|---|
evaluation_id |
str |
required | Evaluation id |
poll_interval |
float |
3.0 |
Seconds between status checks |
timeout |
float |
1800.0 |
Seconds before PollTimeoutError |
sleep |
Callable | None |
None |
Override the sleep function, for tests |
evaluation = client.model_evaluations.wait_for_completion(
evaluation_id="9d8b3c17-2f45-4e6a-b810-5c7e2a94f036",
poll_interval=5.0,
timeout=3600.0,
)
print(evaluation.overall_metrics.f1)
# Poll until status is terminal.
curl -s "https://api.pictograph.io/api/v1/developer/model-evaluations/<evaluation-id>" \
-H "X-API-Key: $PICTOGRAPH_API_KEY"
Returns ModelEvaluation
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
A timeout does not stop the run. The evaluation keeps going server-side and can be
read later with get.
list
The organization’s evaluations, newest first, optionally narrowed to one model or one dataset.
Source: ModelEvaluations.list
| Arg | Type | Default | Notes |
|---|---|---|---|
model |
str | None |
None |
Only this model’s evaluations, by name or id |
dataset_name |
str | None |
None |
Only this dataset’s evaluations |
limit |
int |
50 |
Page size, maximum 200 |
history = client.model_evaluations.list(
model="Swift Falcon",
limit=20,
)
for run in history:
score = run.overall_metrics.f1 if run.overall_metrics else None
print(run.id, run.status, score)
curl -s "https://api.pictograph.io/api/v1/developer/model-evaluations?model_id=6f1c2f0e-6a1e-4a55-9f3a-2f2d3b4c5d6e&limit=20" \
-H "X-API-Key: $PICTOGRAPH_API_KEY"
Returns list[ModelEvaluation]
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
The REST filters take ids: model_id, and project_id for the dataset. This
endpoint returns a single page and takes no offset, so raise limit rather than
paging.
cancel
Stop a pending or running evaluation. First terminal status wins, so cancelling
an already-finished run is a 409 rather than a silent no-op.
Source: ModelEvaluations.cancel
| Arg | Type | Default | Notes |
|---|---|---|---|
evaluation_id |
str |
required | Evaluation id |
evaluation = client.model_evaluations.cancel(
evaluation_id="9d8b3c17-2f45-4e6a-b810-5c7e2a94f036",
)
print(evaluation.status)
curl -s -X POST "https://api.pictograph.io/api/v1/developer/model-evaluations/9d8b3c17-2f45-4e6a-b810-5c7e2a94f036/cancel" \
-H "X-API-Key: $PICTOGRAPH_API_KEY"
Returns ModelEvaluation
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
iter
Auto-paging iterator over the org’s evaluations.
Source: ModelEvaluations.iter
| Arg | Type | Default | Notes |
|---|---|---|---|
model |
str | None |
None |
Model name or id. |
dataset_name |
str | None |
None |
Dataset name. |
page_size |
int |
50 |
Rows fetched per underlying request. Tuning only - the iterator yields every item either way. |
for evaluation in client.model_evaluations.iter(model="Swift Falcon"):
score = evaluation.overall_metrics.f1 if evaluation.overall_metrics else None
print(evaluation.id, evaluation.status, score)
# The endpoint pages with limit + offset; `iter` walks it for you.
curl -s "https://api.pictograph.io/api/v1/developer/model-evaluations?limit=50&offset=50" \
-H "X-API-Key: $PICTOGRAPH_API_KEY"
Returns OffsetPager[ModelEvaluation]
ModelEvaluation · 22 fields
class ModelEvaluation(BaseModel):
"""A model-evaluation run + its metric summary."""
id: str
organization_id: str
model_id: str
dataset_id: str
export_id: str | None = None
status: Literal['pending', 'running', 'completed', 'failed', 'cancelled']
progress: int = 0
iou_threshold: float = 0.5
confidence_threshold: float = 0.5
total_images: int = 0
evaluated_images: int = 0
failed_images: int = 0
overall_metrics: EvalOverallMetrics | None = None
per_class_metrics: list[EvalClassMetrics] | None = None
confusion_matrix: EvalConfusionMatrix | None = None
worst_images: list[EvalWorstImage] | None = None
config: dict[str, Any] | None = None
error_message: str | None = None
created_at: datetime | None = None
updated_at: datetime | None = None
started_at: datetime | None = None
completed_at: datetime | None = None
Common errors
| Status | Exception | Cause |
|---|---|---|
| 400 | ValidationError |
Model is not ready, is not a detection or instance-segmentation model, has no weights, or the export is not completed or has no recorded image set |
| 402 | PaymentRequiredError |
Not enough compute credits to start the run |
| 404 | NotFoundError |
Model, export, dataset, or evaluation not found in your organization |
| 409 | ConflictError |
cancel on an evaluation that already finished |