Ru
ГлавнаяКаналыКриптовалюты → Ontology Official Announcement

Ontology Official Announcement

@ontologyannouncements · Криптовалюты
2 980подписчиков сейчас
Подписаться в Telegram

Публикации всего: 49

Постов на странице: 10 30 50 Страница 1 из 5
📌 Ontology MainNet upgrade to v3.1.2 Ontology will perform a scheduled MainNet upgrade to v3.1.2 at block height 20,800,000. What is in it 🔹 PUSH0 (EIP-3855) places zero on the stack in one byte and 2 gas, instead of two bytes and 3 gas. 🔹 BASEFEE (EIP-3198) lets a contract read the current base fee directly on-chain for 2 gas, removing the need for an external data source. 🔹 MCOPY (EIP-5656) copies memory in a single instruction. Copying 256 bytes drops from at least 96 gas to 27, which matters for encoding, byte handling and cryptographic work. 🔹 Transient storage, TSTORE and TLOAD (EIP-1153) adds low-cost state that lasts for one transaction only, at 100 gas each. Node operators All node operators should upgrade to v3.1.2 as soon as possible and before block 20,800,000 is reached. Release: https://github.com/ontio/ontology/releases/tag/v3.1.2 Full breakdown 👉 https://ont.io/news/ontology-mainnet-v3-1-2-four-ethereum-opcodes-arrive-on-the-ontology-evm/
Перейти к публикации →
📌 Ontology at Eight: verified human data for the AI economy Eight years ago the Ontology MainNet went live, and it has run without interruption ever since. But our eighth anniversary is not really about looking back. It is about what all of it was building toward. AI is only as good as the data it learns from, and the industry is moving away from scraped content toward high-quality human data: consent-based, and provably created by a real person. The problem is supply. The people who create data rarely share in its value, while big tech has earned more than $1.3 trillion from user-generated data. That is the gap we have spent eight years preparing to fill. ONTO Wallet stays a multi-chain Web3 wallet and is now adding an identity and verified human data platform: you own the data you create, build a verified profile, and earn rewards by contributing it on your own terms. Read the full piece 👉 https://ont.io/news/ontology-verified-human-data/
Перейти к публикации →
📌 When human oversight becomes a compliance requirement "We had humans in the loop" is a description of a process. "Prove it" is a demand for evidence, and evidence has properties good intentions do not. It has to name specific people, show they were distinct real humans rather than sybils or one-shot contractors, show their judgement held up over time, and make every contribution attributable, timestamped and tamper-evident. The ground is already moving. The EU AI Act requires human oversight for high-risk AI under Article 14, and the frontier labs themselves are warning that recursive self-improvement could quietly drop the human from the loop. The moment oversight has to satisfy an auditor, attestation ("trust us, qualified humans reviewed this") is not enough. You need provenance: a record someone who does not trust you can check. Run the self-check 👉 https://ont.io/news/human-oversight-documentation/
Перейти к публикации →
📌 New: when the judge shares the blind spot Roll two fair dice. You are told at least one is a six. The probability that both are sixes is not 1 in 6, it is 1 in 11: the clue removes every outcome with no six, leaving 11 equally likely cases, one of which is the double six. The fast answer assumes an independence the clue already broke. Avena et al. tested eight state-of-the-art models on problems built exactly to trigger that shortcut. On the counterintuitive items the models failed consistently and predictably, and chain-of-thought did not reliably rescue them. The trouble is that those same models now do the grading. When an LLM-as-judge carries a reasoning blind spot, a reward model trained on its preferences inherits it, and model-on-model agreement certifies consensus rather than correctness. The only check that does not share the failure mode is human ground truth you can verify: evaluators whose reasoning consistency is measured and tracked over time, carried on a stable identity with signed contributions. W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. ONT ID and ONTO Wallet are the substrate. Day 1 of Ontology Roundup, Issue 04. Try the dice trap 👉 https://ont.io/news/llm-as-judge-blind-spots/
Перейти к публикации →
🎲 We're running a little experiment today, and we want you in it. One dice question. One trap almost everyone falls for, the AI models included. The poll is live on X right now 👉 https://x.com/OntologyNetwork/status/2066416847499485589?s=20 Vote, argue it out in the replies, show your working, and tag a friend who reckons they're good at probability. The more wrong answers, the better the point we're making. No Googling. No AI. That is cheating, and it gets this one wrong anyway. Reveal and the full piece this afternoon. 👀
Перейти к публикации →
A checklist today, because the SFT-vs-RL argument is eating attention that belongs one layer up. Whichever recipe wins for reasoning models, both consume step-level human evaluation, and almost nobody can defend theirs. Five questions sort the defensible pipelines from the rest: who made each judgement, does it carry its rubric, would you notice a drifting evaluator, can experts prove credentials without exposing identity, and does revocation actually propagate. Fewer than three yes answers means the pipeline, not the recipe, is the binding constraint.
Перейти к публикации →
📌 New: Evaluator-backed benchmarking, after the MLE-Bench moment MLE-Bench has been quietly contested across r/MachineLearning over the last week. The skepticism is not really about any single metric inside the benchmark; it is about whether a static benchmark structure can survive sustained adversarial attention from teams with economic incentive to game it. The standard answer (better methodology, rotating held-out sets, broader task coverage) is real and partial. None of it fixes the structural problem: the benchmark as an artefact is a fixed target. Evaluator-backed benchmarking is the structural counter. Every judgement contributing to a published benchmark statistic traces back to a stable evaluator identity (a W3C DID the evaluator controls), a signed verifiable credential carrying the rubric version and expertise attestations, longitudinal consistency credentials, and a status trail for revocations. The benchmark stops being a number the publisher asks the field to trust. It becomes an artefact any third party can audit at the judgement layer. Issue 02 Monday made this argument at the policy-and-research level with the METR teardown. MLE-Bench is the same warning shot moved one level closer to the user-facing capability claim. The first publishers to ship evaluator-backed benchmarking will be the ones whose results survive the next round of teardowns. This is Day 2 of Ontology Roundup, Issue 03. Read it 👉 https://ont.io/news/evaluator-backed-benchmarking/
Перейти к публикации →
📌 New: Reward models need reward-model QA The recent LongTraceRL work made one thing unavoidable: sparse outcome signals are not enough at the reasoning-trace layer. The field has to evaluate intermediate reasoning steps. Step-level evaluation is a substantially different operation than outcome evaluation: judgements are finer-grained, cognitive load on the evaluator is higher, and the noise floor on any individual rating is correspondingly worse. A reward model trained on step-level data is more sensitive to evaluator quality than the outcome-level reward models the field is used to. Sloppy step-level judgement does not just add noise; it miscalibrates the reward model in structured ways the team training the model may not be measuring. Reward-model QA is the missing layer that turns step-level preference data into trustable training signal. The standards stack is the same one Issue 02 set out: W3C Decentralized Identifiers anchor stable evaluator identity, W3C Verifiable Credentials carry signed step-level contributions with rubric versioning, W3C Bitstring Status Lists handle revocation. With those in place, the reward-model team can defend who made each judgement, under what methodology, with what calibration history. This is Day 1 of Ontology Roundup, Issue 03. Read it 👉 https://ont.io/news/reward-model-qa-longtracerl/
Перейти к публикации →
📌 New: The evaluator uniqueness primitive: from sybil resistance to agent evaluation Closing piece for Issue 02. The week opened with the METR teardown and traced the credibility-event pattern through preference data integrity (Tuesday) and longitudinal evaluation (Wednesday). This piece folds together the two threads still open: chronic sybil contamination in preference-data marketplaces, and the agent decision evaluation vacuum that has not yet crystallised into a named problem. Both are solved by the same primitive. Selective disclosure (W3C Verifiable Credentials 2.0 + IETF RFC 9901 SD-JWT) lets a credentialed issuer attest that an evaluator is one unique person, certified by a trust framework, without disclosing identity or demographics. The reward-model team gets the uniqueness guarantee. The evaluator gets privacy. Neither has to compromise. The same mechanic becomes the ground-truth layer for agent decision evaluation when that problem crystallises later this year. The primitive that closes all five Issue 02 threads (benchmark provenance, preference data integrity, longitudinal evaluation, sybil resistance, agent decision evaluation) is human judgement with verifiable uniqueness. The standards work has been done. ONT ID and ONTO Wallet are the substrate. This is Day 5 of Ontology Roundup, Issue 02. Closing piece. Read it 👉 https://ont.io/news/evaluator-uniqueness-closer/
Перейти к публикации →
📌 New: Continuous training needs continuous evaluators Deployed models do not sit still anymore. Retrained, fine-tuned, instruction-extended, behaviourally patched on a cadence measured in weeks, sometimes in days. Last week's Prism paper (Tang et al., arXiv 2605.26110) treats multimodal continual instruction tuning as the deployed reality and flags that the field is hindered by severe engineering bottlenecks. The bottlenecks on the model side are well-defined. The bottlenecks on the evaluation side are larger and quieter. A snapshot evaluator pool against a continually retrained model is the slow version of a contaminated reward dataset. The published delta between version N and N+1 is the sum of two things: actual model behaviour change and cohort composition change. Most teams cannot separate the two. The cost surfaces months later as benchmarks that no longer agree with one another and methodology questions that cannot be resolved without going back to data the pipeline did not keep. Longitudinal evaluation is the property that the evaluator cohort is observable over time, the same way the model is. Stable evaluator identity across batches, signed and timestamped contributions, auditable cohort composition. W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. The standards have been mature for years. This is Day 3 of Ontology Roundup, Issue 02. Read it 👉 https://ont.io/news/longitudinal-evaluation/
Перейти к публикации →
1 2 3 4 5

Другие каналы категории