Commits · 3af93a39d7dddadc13fba978da113aa847509ee5 · SCALE / Code / external / pugixml

Oct 21, 2017

Clarify a note about compact hash behavior during move · 3af93a39

After move some nodes in the hash table can have keys that point to
other; this makes the table somewhat larger but this does not impact
correctness.

The reason is that for us to access a key in the hash table, there
should be a compact_pointer/string object with the state indicating that
it is stored in a hash table, and with the address matching the key. For
this to happen, we had to have put this object into this state which
would mean that we'd overwrite the hash entry with the new, correct
value.

When nodes/pages are being removed, we do not clean up keys from the
hash table - it's safe for the same reason, and thus move doesn't
introduce additional contracts here.

3af93a39

Sep 26, 2017

Fix -Wshadow warning · febf25d1
Arseny Kapoulkine authored 7 years ago

febf25d1

Implement move support for xml_document · a567f12d

Arseny Kapoulkine authored 7 years ago

This change implements the initial version of move construction and
assignment support for documents.

When moving a document to another document, we always make sure move
target is in "clean" state (empty document), and proceed by relocating
all structures in the most efficient way possible.

Complications arise from the fact that the root (document) node is
embedded into xml_document object, so all pointers to it have to change;
this includes parent pointers of all first-level children as well as
allocator pointers in all memory pages and previous pointer in the first
on-heap memory page.

Additionally, compact mode makes everything even more complicated
because some of the pointers we need to update are stored in the hash
table (in fact, document first_child pointer is very likely to be there;
some parent pointers in first-level children will be using
compact_shared_parent but some won't be) which requires allocating a new
hash table which can fail.

Some details of this process are not fully fleshed out, especially for
compact mode; and this definitely requires many tests.

a567f12d

Jul 18, 2017

Fix Clang/C2 compatibility · 77d7e603

Arseny Kapoulkine authored 7 years ago

Clang/C2 does not implement __builtin_expect; additionally we need to
work around deprecation warnings for fopen by disabling them.

77d7e603

Jun 23, 2017

Use PUGI__MSVC_CRT_VERSION instead of _MSC_VER · 853333cd

Arseny Kapoulkine authored 7 years ago

It's not clear whether we still need PUGI__MSVC_CRT_VERSION, but it's
more consistent for now to use it for _snprintf_s since this is relying
on a CRT extension, not on a compiler feature.

853333cd

Jun 22, 2017

Deprecate xml_document::load(const char*) and xml_node::select_single_node · 2252927c

Arseny Kapoulkine authored 7 years ago

These functions were deprecated via comments in 1.5 but never got the
deprecated attribute; now is the time!

Using deprecated functions produces a warning; to silence it, this
change moves the relevant tests to a separate translation unit that has
deprecation disabled.

2252927c

Jun 19, 2017

Change PUGI__SNPRINTF to use _countof for MSVC · 208e2cf0

Arseny Kapoulkine authored 7 years ago

The macro only works correctly when the input argument is an array with
a statically known size - pointers or arrays decayed to pointers won't
work silently.

While this is unlikely to surface issues that aren't caught in
tests/code review, use _countof for MSVC to prevent such code from
compiling.

208e2cf0

Jun 16, 2017

Fix BorlandC compilation · b6995f06
Arseny Kapoulkine authored 7 years ago
```
Rename partition to partition3 to resolve conflicts with std::partition.
```
b6995f06

Refactor snprintf support · 95f013ba

Arseny Kapoulkine authored 7 years ago

Instead of branching code at each invocation site, use variadic macros
to create a wrapping macro that use snprintf for the buffer of a
statically known size.

Variadic macros are supported by all C++11 compilers, as is snprintf;
on MSVC 2005+ we don't necessarily have snprintf, but we can use
_snprintf_s with _TRUNCATE to get the same behavior. In all other cases
we fall back to sprintf, that (theoretically) can lead to a stack buffer
overflow.

In practice all snprintfs used in pugixml use buffers that should be
large enough to never be overflown but snprintf is safe even if this is
not the case.

95f013ba

Use buffer with a static size in convert_number_to_mantissa_exponent · 207bc788

Arseny Kapoulkine authored 7 years ago

We use references to arrays elsewhere in the codebase and there's just
one caller for this function so it's easier to fix the size.

This will simplify snprintf refactoring.

207bc788

Jun 15, 2017
- Mark all assert(false) statements as unreachable · b3b44841
  Arseny Kapoulkine authored 7 years ago
  
  Now we can exclude these from code coverage since it's logically impossible to hit them in tests.
  b3b44841
Jun 11, 2017
- use snprintf if available, _snprintf or sprintf otherwise · 0d8022ec
  Renaud Guillard authored 7 years ago
  
  0d8022ec
Jun 05, 2017
- use _snprintf if MSVC · 810f1f60
  Renaud Guillard authored 7 years ago
  
  810f1f60
Jun 04, 2017
- use snprintf instead of sprintf · b5e9d933
  Renaud Guillard authored 7 years ago
  
  b5e9d933
Apr 04, 2017

Work around -fsanitize=integer issues · 38edf255

Arseny Kapoulkine authored 7 years ago

Integer sanitizer is flagging unsigned integer overflow in several
functions in pugixml; unsigned integer overflow is well defined but it
may not necessarily be intended.

Apart from hash functions, both string_to_integer and integer_to_string
use unsigned overflow - string_to_integer uses it to perform
two-complement negation so that the bulk of the operation can run using
unsigned integers. This makes it possible to simplify overflow checking.
Similarly integer_to_string negates the number before generating a
decimal representation, but negating is impossible without unsigned
overflow or special-casing certain integer limits.

For now just silence the integer overflow using a special attribute;
also move unsigned overflow into string_to_integer from get_value_* so
that we have fewer functions marked with the attribute.

Fixes #133.

38edf255

Mar 22, 2017
- Add missing PUGI__FN to string_to_integer · 101f3288
  Arseny Kapoulkine authored 8 years ago
  
  101f3288
- Revert "Fix gcc-4.8 compilation warning when using -Wstrict-overflow" · 956be4ca
  Arseny Kapoulkine authored 8 years ago
  
  This reverts commit 79109a85. This warning does not happen on gcc-4.8.4; the workaround introduces an unsigned integer overflow which results in a runtime error when compiled with integer sanitizer.
  956be4ca
Mar 05, 2017

Silence g++ 7.0.1 -Wimplicit-fallthrough warnings · 87fc170c

Stephan Beyer authored 8 years ago

This is accomplished by putting a // fallthrough
comment at the right place.
This seems to be more portable than an attribute-based
solution like [[fallthrough]] or __attribute__((fallthrough)).

87fc170c

Mar 03, 2017

Simplify compact_hash_table implementation · 8ce4592e

Arseny Kapoulkine authored 8 years ago

Instead of a separate implementation for find/insert, use just one that
can do both. This reduces the code size and simplifies code coverage;
the resulting code is close to what we had in terms of performance and
since hash table is a fall back should not affect any real workloads.

8ce4592e

Feb 09, 2017

Add invalid type assertion for offset_debug · d4c456bd

Arseny Kapoulkine authored 8 years ago

This will make sure we don't forget to implement offset_debug for new
node types if they ever happen (really it's mostly for consistency).

d4c456bd

Feb 08, 2017

Add invalid type assertion for offset_debug · 0991c1d2

Arseny Kapoulkine authored 8 years ago

This will make sure we don't forget to implement offset_debug for new
node types if they ever happen (really it's mostly for consistency).

0991c1d2

Feb 07, 2017

XPath: Simplify sorting implementation · 2162a0d8

Arseny Kapoulkine authored 8 years ago

Instead of a complicated partitioning scheme that tries to maintain the
equal area in the middle, use a scheme where we keep the equal area in
the left part of the array and then move it to the middle.

Since generally sorted arrays don't contain many duplicates this extra
copy is not too expensive, and it significantly simplifies the logic and
maintains good complexity for sorting arrays with many equal elements
nonetheless (unlike Hoare partitioning).

Instead of a median of 9 just use a median of 3 - it performs pretty
much identically on some internal performance tests, despite having a
bit more comparisons in some cases.

Finally, change the insertion sort threshold to 16 elements since that
appears to have slightly better performance.

2162a0d8

XPath: Optimize insertion_sort · 774d5fe9

Arseny Kapoulkine authored 8 years ago

The previous implementation opted for doing two comparisons per element
in the sorted case in order to remove one iterator bounds check per
moved element when we actually need to copy. In our case however the
comparator is pretty expensive (except for remove_duplicates which is
fast as it is) so an extra object comparison hurts much more than an
iterator comparison saves.

This makes sorting by document order up to 3% faster for random
sequences.

774d5fe9

Feb 06, 2017
- XPath: Remove redundant calls from xml_node::select_nodes et al · 8cc3144e
  Arseny Kapoulkine authored 8 years ago
  
  Instead of delegating to a method that just forwards the call to xpath_query call the relevant method directly.
  8cc3144e
- XPath: Remove evaluate_string_impl · 00e39c58
  Arseny Kapoulkine authored 8 years ago
  
  It adds one stack frame to string query evaluation and does not really simplify the code.
  00e39c58
Feb 04, 2017

XPath: Simplify evaluation error flow · bcc7ed57

Arseny Kapoulkine authored 8 years ago

Instead of having two checks for out-of-memory when exceptions are
enabled, do just one and decide what to do based on whether we can
throw.

bcc7ed57

Feb 03, 2017

XPath: Clean up out-of-memory parse error handling · 33159924

Arseny Kapoulkine authored 8 years ago

Instead of relying on a specific string in the parse result, use
allocator error state to report the error and then convert it to a
string if necessary.

We currently have to manually trigger the OOM error in two places
because we use global allocator in rare cases; we don't really need to
do this so this will be cleaned up later.

33159924

Feb 02, 2017

Remove redundant branch from xml_node::path() · 0e3ccc73

Arseny Kapoulkine authored 8 years ago

The code works fine regardless of the *j->name check, and omitting this
makes the code more symmetric between the "count" and "write" stage;
additionally this improves coverage - due to how strcpy_insitu works
it's not really possible to get an empty non-NULL name in the node.

0e3ccc73

Jan 31, 2017

Remove null pointer test from first_element_by_path · 9c7897b8
Arseny Kapoulkine authored 8 years ago
```
All other functions treat null pointer inputs as invalid; now this
function does as well.
```
9c7897b8

XPath: Remove (re)allocate_throw and setjmp · f500435c

Arseny Kapoulkine authored 8 years ago

Now error handling in XPath implementation relies on explicit error
propagation and is converted to an appropriate result at the end.

f500435c

XPath: Replace all (re)allocate_throw with (re)allocate_nothrow · 9e40c585

Arseny Kapoulkine authored 8 years ago

This generates some out-of-memory code paths that are not covered by
existing tests, which will need to be resolved later.

9e40c585

XPath: Fix reallocate_nothrow to preserve existing state · c370d119

Arseny Kapoulkine authored 8 years ago

Instead of rolling back the allocation and trying to allocate again,
explicitly handle inplace reallocate if possible, and allocate a new
block otherwise.

This is going to be important once we use reallocate_nothrow from a
non-throwing context.

c370d119

XPath: Use nonthrowing allocations in duplicate_string · 1a2e4b88
Arseny Kapoulkine authored 8 years ago
```
This requires explicit error handling for xpath_string::data calls.
```
1a2e4b88

XPath: Throw std::bad_alloc if we got an out-of-memory error · ac150d50

Arseny Kapoulkine authored 8 years ago

This allows us to gradually convert exception handling of out-of-memory
during evaluation to a non-throwing approach without changing the
observable behavior.

ac150d50

Jan 30, 2017

XPath: Reword brace mismatch errors for clarity · 1b3e8614
Arseny Kapoulkine authored 8 years ago

1b3e8614

XPath: Improve error message for expressions like .[1] · 1ed6d210

Arseny Kapoulkine authored 8 years ago

W3C specification does not allow predicates after abbreviated steps.
Currently this results in parsing terminating at the step, which leads
to confusing error messages like "Invalid query" or "Unmatched braces".

1ed6d210

XPath: Track allocation errors more explicitly · bc1e4446

Arseny Kapoulkine authored 8 years ago

Any time an allocation fails xpath_allocator can set an externally
provided bool. The plan is to keep this bool up until evaluation ends,
so that we can use it to discard the potentially malformed result.

bc1e4446

XPath: Provide non-throwing and throwing allocations in xpath_allocator · 635fe028

Arseny Kapoulkine authored 8 years ago

For both allocate and reallocate, provide both _nothrow and _throw
functions; this change renames allocate() to allocate_throw() (same for
reallocate) to make it easier to change the code to remove throwing
variants.

635fe028

XPath: Minor error handling refactoring · 6abf1d7c
Arseny Kapoulkine authored 8 years ago
```
Handle node type error before creating expression node
```
6abf1d7c

XPath: Route out-of-memory errors through the exceptionless path · 4fa2241d

Arseny Kapoulkine authored 8 years ago

We currently need to convert error based on the text to a different type
of C++ exceptions when C++ exceptions are enabled.

4fa2241d