pugixml

mirror of https://github.com/zeux/pugixml.git synced 2025-01-14 01:47:55 +08:00

Author	SHA1	Message	Date
Arseny Kapoulkine	7c6d0010b3	Merge pull request #170 from zeux/move This change implements move ctor and assign support for xml_document. All node handles remain valid after the move and point to the new document; the only exception is the document node itself (that remains unmoved). Move is O(document size) in theory because it needs to relocate immediate document children (there is just one in conformant documents) and all memory pages; in practice the memory pages only need the header adjusted, which is ~0.1% of the actual data size. Move requires no allocations in general, except when using compact mode where some moves need to grow the hash table which can fail (throw). Fixes #104	2017-11-13 13:24:43 -08:00
Arseny Kapoulkine	3860b5076f	Fix -Wshadow warning	2017-11-13 09:27:38 -08:00
Arseny Kapoulkine	58611c8702	tests: Add compact move tests This helps make sure our error handling logic works and is exercised.	2017-11-13 08:59:16 -08:00
Arseny Kapoulkine	4bd8771c2f	Implement correct move error handling for compact mode In compact mode, we currently can not support zero-allocation moves since some pointer assignments required during the move need to allocate hash table slots. This is mostly applicable to xml_document_struct::first_child, since the pointer to this element is used as a hash table key, but there are some contrived cases where parents of root's children need a hash slot and didn't have it before. These cases can be fixed by changing the compact encoding to be a bit more move friendly, but for now it's easier to handle the error and throw/return during move. When this happens, the source document doesn't change.	2017-11-13 08:57:16 -08:00
Arseny Kapoulkine	91a3c28862	Add count argument to compact_hash_table::rehash/reserve This allows us to do a single reserve for a known amount of assignments that is larger than the default minimum per reserve (16).	2017-11-13 08:37:34 -08:00
Arseny Kapoulkine	6016e2180e	CMake: Add __declspec(dllexport) for shared library builds This makes sure that MSVC shared library build actually exports all the needed symbols and generates import table. Somehow, this is actually enough to make pugixml link as a DLL - there's no need to specify __declspec(dllimport) even though pugixml exports classes via DLL. Fixes #113.	2017-11-12 20:21:46 -08:00
Arseny Kapoulkine	492ebc22bc	tests: Fix expansion-to-defined warning This warning is new as of GCC 7 and highlights undefined behavior in the preprocessor that ASAN detection was relying on.	2017-11-10 21:35:59 -08:00
Arseny Kapoulkine	6fe31d1477	build: Simplify config=sanitize These days OSX clang supports UB sanitizer so we can just use the same settings for all systems.	2017-10-29 21:24:04 -07:00
Arseny Kapoulkine	ba9504325e	build: Switch fuzz builds to use Clang 5.0 sanitize=fuzzer The old fuzzer location is deprecated; this also makes it almost trivial to fuzz, provided that the clang is set up correctly... on Ubuntu 17.10, a command sequence like this works now: sudo apt install clang-5.0 sudo apt install libfuzzer-5.0 sudo cp /usr/lib/llvm-5.0/lib/libFuzzer.a /usr/lib/libLLVMFuzzer.a CXX=clang++-5.0 make fuzz_parse	2017-10-29 19:54:48 -07:00
Arseny Kapoulkine	0504fa4e90	tests: Add more tests for document move These tests currently fail for compact mode because of ->reserve() failing.	2017-10-26 08:36:05 -07:00
Arseny Kapoulkine	3af93a39d7	Clarify a note about compact hash behavior during move After move some nodes in the hash table can have keys that point to other; this makes the table somewhat larger but this does not impact correctness. The reason is that for us to access a key in the hash table, there should be a compact_pointer/string object with the state indicating that it is stored in a hash table, and with the address matching the key. For this to happen, we had to have put this object into this state which would mean that we'd overwrite the hash entry with the new, correct value. When nodes/pages are being removed, we do not clean up keys from the hash table - it's safe for the same reason, and thus move doesn't introduce additional contracts here.	2017-10-20 21:57:14 -07:00
Arseny Kapoulkine	b0fc587a7f	tests: Add more move tests We now check that appending a child to a moved document performs no allocations - this is already the case, but if we neglected to copy the allocator state this test would fail.	2017-10-20 21:53:42 -07:00
Arseny Kapoulkine	50bc0d5a69	tests: Adjust move coverage tests Large test wasn't testing shared parent condition properly - add one more level of hierarchy so that it works as expected.	2017-09-25 22:54:42 -07:00
Arseny Kapoulkine	26ead385a7	tests: Add more move tests Add a test that checks that static buffer pointer was moved correctly by checking if offset_debug still works.	2017-09-25 22:52:26 -07:00
Arseny Kapoulkine	402b967fa9	tests: Add more move tests Make sure we have coverage for empty documents and for large documents that trigger compact_shared_parent != root for some pages.	2017-09-25 22:47:10 -07:00
Arseny Kapoulkine	faba4786c0	tests: Add more document move tests Verify that move doesn't allocate and that it preserves structures required for tree memory management and append_buffer in tact.	2017-09-25 22:38:30 -07:00
Arseny Kapoulkine	febf25d1af	Fix -Wshadow warning	2017-09-25 21:48:37 -07:00
Arseny Kapoulkine	6eb7519dba	tests: Add basic move tests These just verify that move ctor/assignment operator work as expected in simple cases - there are a number of ways in which the internal structure can be incorrect...	2017-09-25 21:37:56 -07:00
Arseny Kapoulkine	a567f12d76	Implement move support for xml_document This change implements the initial version of move construction and assignment support for documents. When moving a document to another document, we always make sure move target is in "clean" state (empty document), and proceed by relocating all structures in the most efficient way possible. Complications arise from the fact that the root (document) node is embedded into xml_document object, so all pointers to it have to change; this includes parent pointers of all first-level children as well as allocator pointers in all memory pages and previous pointer in the first on-heap memory page. Additionally, compact mode makes everything even more complicated because some of the pointers we need to update are stored in the hash table (in fact, document first_child pointer is very likely to be there; some parent pointers in first-level children will be using compact_shared_parent but some won't be) which requires allocating a new hash table which can fail. Some details of this process are not fully fleshed out, especially for compact mode; and this definitely requires many tests.	2017-09-25 19:31:18 -07:00
Arseny Kapoulkine	a569e6a737	Switch to sudo=false Travis env	2017-09-24 22:46:48 -07:00
Arseny Kapoulkine	900a1cc943	docs: Clarify Unicode validation behavior It has always been the case that pugixml does not perform Unicode validation or name/tag Unicode character class validation, but it wasn't very obvious from documentation. Fixes #162	2017-08-29 20:46:30 -07:00
Arseny Kapoulkine	4f2ad720c8	docs: Update encoding conversion description We support Latin-1 and automatically detect it by parsing the encoding from document declaration; both of these were omitted from the description of the automatic detection. Additionally, the description has been rewritten to be more concise and a bit more abstract - there's no need to specify the algorithm precisely here. Fixes #158.	2017-08-21 20:54:38 -07:00
Arseny Kapoulkine	50952c0a5e	scripts: Fix NuGet VS2017 build Due to a typo in build script v141 binaries were built using VS2015 instead of VS2017. Fixes #157.	2017-08-20 07:12:09 +01:00
Arseny Kapoulkine	f423cec11b	scripts: Disable LTCG for VS2017 Using LTCG restricts the resulting .lib files to a specific compiler version, causing version conflicts when the compiler gets updated without changing the toolset version. VS2017 now has two incompatible compilers, 15.0 and 15.3, both of which use toolset v141...	2017-08-18 21:44:47 +01:00
Arseny Kapoulkine	77d7e60379	Fix Clang/C2 compatibility Clang/C2 does not implement __builtin_expect; additionally we need to work around deprecation warnings for fopen by disabling them.	2017-07-17 22:15:35 -07:00
Arseny Kapoulkine	ed86ef32b3	Update README.md Switch codecov.io URLs to https	2017-06-29 01:03:55 -07:00
Arseny Kapoulkine	cfa64676be	Update README.md	2017-06-29 00:53:59 -07:00
Arseny Kapoulkine	f3e0f4249c	tests: Add more stream coverage tests These new tests test that tellg() can fail when being called the second time, which leads to seekable implementation failing.	2017-06-23 08:44:52 -07:00
Arseny Kapoulkine	4564d31c76	tests: Add stream coverage tests These tests simulate various error conditions when reading data from streams - seeks failing in seekable streams, underflow throwing an exception causing read to set badbit, etc. This change also adjusts memory thresholds to cause a reliable out of memory during construction of a final buffer for non-seekable streams.	2017-06-23 07:48:09 -07:00
Arseny Kapoulkine	20a8eced3b	tests: Fix PUGIXML_WCHAR_MODE build	2017-06-22 22:18:16 -07:00
Arseny Kapoulkine	3870217381	tests: Add more XPath out of memory tests This fixes missing coverage in translate_table_generate and xpath_node_set_raw::append.	2017-06-22 22:11:43 -07:00
Arseny Kapoulkine	5867aff943	tests: Make using namespace more explicit Hiding using namespace in common.hpp is somewhat surprising so remove common.hpp and move using namespace into all .cpp files that need it.	2017-06-22 20:41:08 -07:00
Arseny Kapoulkine	4b371e10ee	tests: Remove redundant pugi:: qualifier Most tests have `using namespace pugi` which makes explicit qualifications unnecessary.	2017-06-22 20:33:02 -07:00
Arseny Kapoulkine	853333cd70	Use PUGI__MSVC_CRT_VERSION instead of _MSC_VER It's not clear whether we still need PUGI__MSVC_CRT_VERSION, but it's more consistent for now to use it for _snprintf_s since this is relying on a CRT extension, not on a compiler feature.	2017-06-22 20:28:06 -07:00
Arseny Kapoulkine	2252927c04	Deprecate xml_document::load(const char*) and xml_node::select_single_node These functions were deprecated via comments in 1.5 but never got the deprecated attribute; now is the time! Using deprecated functions produces a warning; to silence it, this change moves the relevant tests to a separate translation unit that has deprecation disabled.	2017-06-22 09:13:10 -07:00
Arseny Kapoulkine	94ef7b3a03	Merge pull request #151 from zeux/nuget Rework NuGet package building	2017-06-20 21:32:11 -07:00
Arseny Kapoulkine	88d43a7ebc	scripts: Refactor nuget_build.ps1 Unify build paths in all MSBuild VS projects and extract common build logic into functions. Note that this change changes both VS2010 and VS2013 projects to have more predictable output paths and fixed output file name (pugixml).	2017-06-20 21:11:35 -07:00
Arseny Kapoulkine	fbc7085c14	scripts: Clarify the linkage settings in package description Also improve linkage description	2017-06-20 21:11:35 -07:00
Arseny Kapoulkine	d2b0328198	Remove CoApp msi installation We build NuGet package manually now so we don't need CoApp.	2017-06-20 21:11:35 -07:00
Arseny Kapoulkine	7d651ac3b6	Update .gitignore	2017-06-20 21:11:35 -07:00
Arseny Kapoulkine	a7c4070df7	scripts: Switch to manual NuGet package with both CRT linkages We'd like to build pugixml with both static & dynamic CRT and put it all in one NuGet package. CoApp sort of allows us to do this via dynamic/static pivots, but it does not let us customize the names of the pivots and additionally has some bugs with the project setup. Their project modifications are also much more complicated - really, at this point we should do this ourselves. Create a simple native NuGet package with Linkage setting that picks the right library, and package all libraries appropriately. Note that we use the unified path syntax to make it simple to just get the right .lib file from the toolset/platform/configuration/linkage combo.	2017-06-20 21:11:35 -07:00
Arseny Kapoulkine	208e2cf043	Change PUGI__SNPRINTF to use _countof for MSVC The macro only works correctly when the input argument is an array with a statically known size - pointers or arrays decayed to pointers won't work silently. While this is unlikely to surface issues that aren't caught in tests/code review, use _countof for MSVC to prevent such code from compiling.	2017-06-19 07:06:47 -07:00
Arseny Kapoulkine	867bd2583b	Merge pull request #150 from zeux/nuget Add VS2017 to AppVeyor test run	2017-06-18 22:34:08 -07:00
Arseny Kapoulkine	9357837d2e	Add VS2017 to AppVeyor test run This requires moving the list of VS versions out of autotest-appveyor.ps1 and into appveyor.yml.	2017-06-18 22:20:13 -07:00
Arseny Kapoulkine	7418bd0d79	scripts: Cleanup nuget_build.ps1 Correctly check for error codes and don't run .bat file since it doesn't work anyway (the variables it sets aren't accessible in PowerShell, and the path to the script doesn't seem to be the same in VS2017).	2017-06-18 21:11:54 -07:00
Arseny Kapoulkine	ade869ea58	Merge pull request #147 from igagis/master VS2017 project + NuGet support	2017-06-18 20:49:11 -07:00
Arseny Kapoulkine	0027b6ac79	tests: Improve XPath coverage Add memory allocation failure test for concact with a very large list and make sure we have every single axis covered with and without a predicate, with and without a previous step.	2017-06-16 22:45:42 -07:00
Arseny Kapoulkine	08f102f14c	tests: Add even more stream coverage tests Apparently only narrow character streams had out of memory coverage - fix that and also split this into a separate test.	2017-06-16 21:38:55 -07:00
Arseny Kapoulkine	86593c0999	tests: Add more stream coverage tests Cover both char and wchar_t stream loading in a single run instead of using pugi::char_t.	2017-06-16 17:08:00 -07:00
Arseny Kapoulkine	3aa2b40354	tests: Add more coverage tests for stream loading Cover more failure cases and simplify the streambuf implementation a bit.	2017-06-16 16:41:08 -07:00
Arseny Kapoulkine	b6995f06b9	Fix BorlandC compilation Rename partition to partition3 to resolve conflicts with std::partition.	2017-06-16 00:32:01 -07:00
Arseny Kapoulkine	bd23216420	tests: Improve XPath test coverage Add more memory allocation failure tests.	2017-06-16 00:29:14 -07:00
Arseny Kapoulkine	a3664ea971	tests: Expand write_flush coverage Adjust the buffer size to be right on the edge of the overflow, make sure we actually output " instead of ".	2017-06-16 00:09:32 -07:00
Arseny Kapoulkine	d2892be902	tests: Add xml_buffered_writer coverage test This test triggers flush() condition for each optimized write() method.	2017-06-15 23:52:56 -07:00
Arseny Kapoulkine	95f013ba80	Refactor snprintf support Instead of branching code at each invocation site, use variadic macros to create a wrapping macro that use snprintf for the buffer of a statically known size. Variadic macros are supported by all C++11 compilers, as is snprintf; on MSVC 2005+ we don't necessarily have snprintf, but we can use _snprintf_s with _TRUNCATE to get the same behavior. In all other cases we fall back to sprintf, that (theoretically) can lead to a stack buffer overflow. In practice all snprintfs used in pugixml use buffers that should be large enough to never be overflown but snprintf is safe even if this is not the case.	2017-06-15 23:35:20 -07:00
Arseny Kapoulkine	207bc788e9	Use buffer with a static size in convert_number_to_mantissa_exponent We use references to arrays elsewhere in the codebase and there's just one caller for this function so it's easier to fix the size. This will simplify snprintf refactoring.	2017-06-15 22:58:46 -07:00
Arseny Kapoulkine	cd2804d3ee	Merge pull request #145 from noresources/snprintf use snprintf instead of sprintf	2017-06-15 21:34:04 -07:00
Arseny Kapoulkine	0698810abb	Merge pull request #149 from zeux/test-path Improve code coverage	2017-06-15 21:17:26 -07:00
Arseny Kapoulkine	927d321d90	Exclude unreachable lines from code coverage codecov.io does not seem to support lcov regex customization; additionally, we can't just replace unreachable with LCOV_LINE_EXCL in gcov file - so we have to patch the ##### indicator (which suggests the line hasn't been hit) with 1. See also https://github.com/codecov/support/issues/144	2017-06-15 20:58:26 -07:00
Arseny Kapoulkine	b3b44841f0	Mark all assert(false) statements as unreachable Now we can exclude these from code coverage since it's logically impossible to hit them in tests.	2017-06-15 09:26:23 -07:00
Arseny Kapoulkine	c40fd364ce	tests: Add tests for loading special files New tests try to load a folder as an XML document, and a device. Both are intended to exercise some otherwise non-hittable error paths in load_file implementation.	2017-06-15 07:23:49 -07:00
Ivan Gagis	b66ca4f326	use 2 images to build on appveyor	2017-06-15 11:49:10 +03:00
Ivan Gagis	4dc1054104	use powershell instead of cmd	2017-06-15 11:32:46 +03:00
Ivan Gagis	e944623780	Appveyor image set to VS2017	2017-06-15 11:17:34 +03:00
Ivan Gagis	042eae4c83	Appveyor image set to VS2017	2017-06-15 11:14:28 +03:00
Ivan Gagis	d8f9148d36	set v141 tools environment for building	2017-06-15 11:11:36 +03:00
Ivan Gagis	3a8073cca2	Appveyor image set to VS2017	2017-06-15 11:05:28 +03:00
Ivan Gagis	c7131b01f9	VS2017 project	2017-06-15 10:59:33 +03:00
Arseny Kapoulkine	0fbc043183	tests: Increase compact_pointer coverage This adds tests that complete branch coverage in compact pointer encoding/decoding code (previously first_attribute was always encoded using compact encoding in the entire test suite).	2017-06-14 23:50:21 -07:00
Arseny Kapoulkine	52da6f71d0	Increase the minimum CMake version to 2.8.12 This is a followup to 198900eff403982f080958459f1ccb45cdefe9a4. target_include_directories was introduced in 2.8.12, thus CMake 2.6 no longer works.	2017-06-14 23:06:32 -07:00
Renaud Guillard	0d8022eced	use snprintf if available, _snprintf or sprintf otherwise	2017-06-11 18:33:28 +02:00
Renaud Guillard	810f1f600d	use _snprintf if MSVC	2017-06-05 13:31:58 +02:00
Renaud Guillard	b5e9d933ad	use snprintf instead of sprintf	2017-06-04 21:10:19 +02:00
Arseny Kapoulkine	38edf255ae	Work around -fsanitize=integer issues Integer sanitizer is flagging unsigned integer overflow in several functions in pugixml; unsigned integer overflow is well defined but it may not necessarily be intended. Apart from hash functions, both string_to_integer and integer_to_string use unsigned overflow - string_to_integer uses it to perform two-complement negation so that the bulk of the operation can run using unsigned integers. This makes it possible to simplify overflow checking. Similarly integer_to_string negates the number before generating a decimal representation, but negating is impossible without unsigned overflow or special-casing certain integer limits. For now just silence the integer overflow using a special attribute; also move unsigned overflow into string_to_integer from get_value_* so that we have fewer functions marked with the attribute. Fixes #133.	2017-04-03 23:35:24 -07:00
Arseny Kapoulkine	24d1a4562b	Move libFuzzer build to Makefile Now the only thing fuzz_setup.sh does is installing new clang; if system clang supports -fsanitize-coverage then fuzz_setup.sh is not required.	2017-04-03 21:09:37 -07:00
Arseny Kapoulkine	0eb1ddb975	tests: Fix fuzz_setup.sh The script only worked if clang folder was already created.	2017-04-03 20:36:33 -07:00
Arseny Kapoulkine	101f32884f	Add missing PUGI__FN to string_to_integer	2017-03-21 22:06:19 -07:00
Arseny Kapoulkine	956be4ca4b	Revert "Fix gcc-4.8 compilation warning when using -Wstrict-overflow" This reverts commit 79109a8546f963d17522d75112cffcfd8cbe35fc. This warning does not happen on gcc-4.8.4; the workaround introduces an unsigned integer overflow which results in a runtime error when compiled with integer sanitizer.	2017-03-21 21:57:16 -07:00
Arseny Kapoulkine	acfe47ba52	tests: Do not use unsigned underflow in test code This triggers a runtime error under integer sanitizer	2017-03-21 21:47:22 -07:00
Arseny Kapoulkine	c29940ca72	tests: Fix invalid buffer size This was triggering an buffer read overflow with asan.	2017-03-21 10:33:20 -07:00
Arseny Kapoulkine	db98a7e28b	Fix path to fuzzing corpus	2017-03-21 10:28:20 -07:00
Arseny Kapoulkine	640c94f90d	Merge pull request #134 from ogdf/explicit-fallthroughs Silence g++ 7.0.1 -Wimplicit-fallthrough warnings	2017-03-06 07:43:40 -08:00
Stephan Beyer	87fc170cdf	Silence g++ 7.0.1 -Wimplicit-fallthrough warnings This is accomplished by putting a // fallthrough comment at the right place. This seems to be more portable than an attribute-based solution like [[fallthrough]] or __attribute__((fallthrough)).	2017-03-05 22:12:10 +01:00
Arseny Kapoulkine	8ce4592e15	Simplify compact_hash_table implementation Instead of a separate implementation for find/insert, use just one that can do both. This reduces the code size and simplifies code coverage; the resulting code is close to what we had in terms of performance and since hash table is a fall back should not affect any real workloads.	2017-03-03 07:11:22 -08:00
Arseny Kapoulkine	03e4b8de92	Merge pull request #132 from zeux/fuzz Improve fuzzing support	2017-02-11 13:51:39 -08:00
Arseny Kapoulkine	ec984370fb	tests: Fix fuzz_setup.sh Make the file executable, fix Windows newlines and fix clang setup.	2017-02-11 13:17:27 -08:00
Arseny Kapoulkine	ea544eb48b	tests: Add fuzzing dictionaries Hopefully this will allow for better fuzzing coverage	2017-02-11 13:17:02 -08:00
Arseny Kapoulkine	8c62fa9121	tests: Add XPath fuzzing Only fuzz the parser for now.	2017-02-09 07:37:38 -08:00
Arseny Kapoulkine	8b15ae8015	tests: Add a script to set up fuzzing tools This downloads a clang build that has support for instrumentation, and also downloads and compiles libFuzzer.a.	2017-02-09 07:37:04 -08:00
Arseny Kapoulkine	00ef791078	fuzz: Use libFuzzer instead of afl-fuzz This allows us to have faster fuzz cycles since the fuzzer is in-process.	2017-02-09 07:36:32 -08:00
Arseny Kapoulkine	e748f435e5	tests: Increase the number of translate calls This should make the test fail on a 32-bit target.	2017-02-09 07:36:32 -08:00
Arseny Kapoulkine	4bab082a27	tests: Fix clang build	2017-02-09 07:36:32 -08:00
Arseny Kapoulkine	ba39838ab5	tests: Add more XPath out of memory tests	2017-02-09 07:36:31 -08:00
Arseny Kapoulkine	d4c456bdef	Add invalid type assertion for offset_debug This will make sure we don't forget to implement offset_debug for new node types if they ever happen (really it's mostly for consistency).	2017-02-09 07:36:31 -08:00
Arseny Kapoulkine	02c599f52b	tests: Increase the number of translate calls This should make the test fail on a 32-bit target.	2017-02-08 01:18:11 -08:00
Arseny Kapoulkine	b98c914053	tests: Fix clang build	2017-02-08 00:33:51 -08:00
Arseny Kapoulkine	1688f44185	tests: Add more XPath out of memory tests	2017-02-08 00:09:32 -08:00
Arseny Kapoulkine	0991c1d283	Add invalid type assertion for offset_debug This will make sure we don't forget to implement offset_debug for new node types if they ever happen (really it's mostly for consistency).	2017-02-07 20:34:49 -08:00
Arseny Kapoulkine	2162a0d80c	XPath: Simplify sorting implementation Instead of a complicated partitioning scheme that tries to maintain the equal area in the middle, use a scheme where we keep the equal area in the left part of the array and then move it to the middle. Since generally sorted arrays don't contain many duplicates this extra copy is not too expensive, and it significantly simplifies the logic and maintains good complexity for sorting arrays with many equal elements nonetheless (unlike Hoare partitioning). Instead of a median of 9 just use a median of 3 - it performs pretty much identically on some internal performance tests, despite having a bit more comparisons in some cases. Finally, change the insertion sort threshold to 16 elements since that appears to have slightly better performance.	2017-02-07 00:05:50 -08:00
Arseny Kapoulkine	774d5fe9df	XPath: Optimize insertion_sort The previous implementation opted for doing two comparisons per element in the sorted case in order to remove one iterator bounds check per moved element when we actually need to copy. In our case however the comparator is pretty expensive (except for remove_duplicates which is fast as it is) so an extra object comparison hurts much more than an iterator comparison saves. This makes sorting by document order up to 3% faster for random sequences.	2017-02-06 19:28:33 -08:00

1 2 3 4 5 ...

1643 Commits