PureTensor · Intelligence, Distilled

An intelligence that works for you, remembers everything, and never leaves your control

PureTensor gives you a working intelligence of your own. It researches the companies and events you care about, writes verified briefs, and keeps a memory of everything it has learned and everything you have told it. Argus is the first thing it does for you. Built by a research lab that owns and operates its NVIDIA Blackwell fleet, and publishes what it measures.

NVIDIA Blackwell · HQ Mountain View, CA · Compute in the UK · DR in Iceland

It works on its own

Give it a question or a company to follow. It researches, checks its sources, and writes the answer while you do something else.

It remembers everything

Every brief, note, and decision you give it stays with it. Ask next month and it picks up where you left off, instead of starting from zero.

It stays yours

Your briefs, your notes, and everything it has learned about your world belong to you. They are never shared, resold, or used to train someone else's model.

Flagship

Argus: verified written intelligence from open sources

Argus reads funding announcements, filings, hiring signals, and public reporting continuously, then writes assessed intelligence briefs on the companies and events you track. It is built for organisations whose questions cannot go to a third-party cloud: every stage of collection, assessment, and writing runs inside a boundary we control. In production, delivering briefs to clients in the AI infrastructure sector since June 2026.

Sample brief, excerptCompany identity withheld. Shape and fields as delivered to clients, September 2026.
Rank 1 of 5

Clinical-stage AI-native biotech, hybrid posture with an owned GPU and Slurm research cluster

Headquarters
Boston, MA, USA
Team size
40 to 70
Funding
Venture, $70M+ cumulative; round label and close date not published
Discovered via
Careers board, surfaced on a Slurm sweep of an on-site infrastructure posting
From the job posting

Experience with HPC schedulers such as Slurm or batch compute for GPU training jobs.

Assessment

A foundation model for cell signalling is a real training workload, and the infrastructure role owns the GPU and Slurm estate that runs it rather than a cloud account. The position is on-site and spans laboratory computers as well as research compute, which is not how a purely rented cluster gets staffed. Production runs on a public cloud, so the opening is the research cluster specifically, not the whole estate.

Every claim above traces to the company's own announcements and its live job posting.

Assessed, Not Aggregated

A multi-model assessment council scores every document in parallel for novelty, impact, and analytical depth. Only verified material reaches a brief.

Entity & Relationship Mapping

Named entity recognition feeds a live knowledge graph. Companies, people, and events resolve into networks rather than headlines.

Full Provenance

Every claim in a brief traces to its sources. Collection, enrichment, and publication run as one auditable pipeline.

The Platform

Built and run on our own stack

Argus runs on infrastructure PureTensor owns: NVIDIA Blackwell compute, 200G RDMA fabric, petascale distributed storage, and Kubernetes orchestration, operated as one system with no third-party cloud in the path. Sentinel, our autonomic operations layer, keeps it alive: continuous triage across every signal, remediation with graded outcomes, and a versioned memory that turns each fix into a permanent immunity. We run our own company on this stack every day.

Compute & Inference

Current-generation NVIDIA Blackwell inference and training on AMD Zen 5 platforms with terabytes of DDR5 system memory.

PureTensor infrastructure topology diagramCompute nodes connect through a high-speed switching fabric to distributed Ceph storage.COMPUTEGPU NODECOMPUTEGPU NODECOMPUTEGPU NODEFABRIC CORERDMA SWITCHINGCEPHPOOLCEPHPOOLCEPHPOOLCEPHPOOLNVIDIA Blackwell Compute400G Spine · 200G RDMA FabricPetascale Ceph Storage

Platform & Storage

Highly available erasure-coded storage pools with dedicated storage fabric. Kubernetes orchestration across the full stack.

AMD Zen 5·NVIDIA Blackwell·NVIDIA Mellanox·200G RDMA·400G Spine·PCIe Gen5 NVMe·Ceph·Kubernetes

Compute operated in the United Kingdom. Disaster recovery in Iceland. No third-party cloud in the path.

Research & Writing

From the Lab

Everything below was measured on the fleet described above, including the failures. We publish what we run.

Sep 8, 2026Research

Same Answers, Different Confidence: 4-Bit Weights Preserve Accuracy and Destroy Calibration

We ran GLM-5.3-Flash in BF16 and in an NVFP4 quantisation on the same eight-GPU Hopper node, where the 4-bit weights are dequantised before arithmetic so that any difference is a property of the stored weights alone. On our internal frontier benchmark every per-dimension accuracy delta was one item, inside the noise: on accuracy the quantisation is a tie. The one dimension that did not tie is calibration. Both checkpoints answered 85.0% of the calibration items correctly, but the BF16 checkpoint's Brier score was 0.070 and the NVFP4 checkpoint's was 0.170, 2.4 times worse and near the 0.1875 random floor. The 4-bit experts keep the answers and flatten the model's stated confidence, so wrong answers arrive with the same tone as right ones. For tool lanes verified downstream the smaller checkpoint is free; for any lane that consumes the model's confidence, a judge, a scorer, an abstention gate, the calibration loss is the price.

Read
Sep 4, 2026Research

Distributed Inference on Workstation Blackwell, Part 4: Cross-Node Tensor Parallelism Over 200 GbE, and the Five Fixes the SM120 Path Needed

Parts 1 to 3 of this series characterised the fabric between two workstation Blackwell nodes and ran models across it by RPC. This part shards every layer across all four RTX PRO 6000 GPUs over 200 GbE with GPUDirect RDMA, which is the only on-premises shape for models that do not fit one node. Nemotron 3 Ultra 550B served three minutes after a cold boot, answered 10 of 10 on a reasoning ladder, and completed a 90-minute soak at 912 of 912 requests with zero errors; GLM-5.3-Flash needed nine attempts and five distinct fixes in the SM120 attention path before it produced a coherent token, and a qualified community image later took it to 691 tokens per second aggregate at 16 concurrent streams. Two failures turned out to be structural to multi-node rather than to any engine: a custom all-reduce prober that deadlocks every rank before a weight loads, and speculative decoding whose data-dependent acceptance length diverges the ranks' collective counts. We also record a public 1,004.9 tokens-per-second headline that measured a locked repeat loop, and two attempts that produced no model-quality evidence and were not scored.

Read
Sep 1, 2026Engineering

Green Gate, Lost Hunks: Nine Ways an Automated Merge Train Dropped Merged Code While Every Test Passed

An adversarial code-review wave produced about 330 pull requests across 21 repositories in one day, every one carrying a semantic-version bump in the same commit, so every merge conflicted on the version surface. The merge train we wrote to restack, restamp, gate, and merge them lost merged work in nine distinct ways, and every one shipped a green test gate: a base-wins rule that dropped a pull request's own middleware from a file that was both a version file and a code file; a version extractor that took an IP address for a version and rewrote it over seven consecutive merges; a silent checkout failure that force-pushed one pull request's content over the next. The only control that caught them was a per-pull-request marker pass on the final default branch, run by a different actor and asking a different question: not whether the tree is healthy, but whether this change arrived.

Read
About

Built From First Principles

What happens when you stop renting intelligence and start building it? We design and operate our own AI infrastructure, from the network fabric to the inference stack, because serious research requires systems you understand completely. Not abstractions on top of abstractions, but hardware you can touch, models you can inspect, and pipelines you control end to end.

What We Believe

Own the stack

From NVIDIA silicon to Ceph storage to Kubernetes orchestration, we operate every layer. No black boxes.

Research in the open

Our findings, benchmarks, and post-mortems are published, including the failures. Science requires scrutiny.

Build what matters

We don’t chase benchmarks. We build systems that solve real problems for real organisations.

Heimir Helgason

Heimir HelgasonLinkedIn ↗

Founder & Chief Architect

Designed and built PureTensor's AI platform from bare metal. 200G RDMA fabric, petascale distributed storage, NVIDIA Blackwell inference. Background in algorithmic trading, cross-border capital markets, and entrepreneurship. Deep expertise in autonomous agent systems, sovereign infrastructure design, and large-scale model deployment.

Ahmed W. Khalil

Ahmed W. KhalilLinkedIn ↗

Strategic Advisor

CFA charterholder with a career spanning top-tier international law, institutional capital allocation, and cross-border deal execution across EMEA. Advises on capital strategy, investor relations, and international market expansion.

We are growing. Reach out.

Contact

Get in Touch

Interested in collaborating, investing, or just talking about AI infrastructure? We'd like to hear from you.

Location

Mountain View, California · London, United Kingdom

Every message is read by the team.